Restatement Recall
The public-algorithm floor for forensic AI on SEC restatements. Reproducible, peer-reviewable, openly methodology-disclosed.
v1.1 result: 100% of eligible SEC restatements caught from the pre-fraud- disclosure 10-K alone, using OloLand's full forensic stack (Beneish M-Score + EBITDA bridge + cross-doc reconciliation + structured risk taxonomy (300+ leaf-level checks) in RiskExtractionMode.FORENSIC).
n=10, 95% CI: 72%–100%, single-seed
Why publish this number?
Big-4 QoE shops don't publish recall numbers. Neither do AI competitors (Hebbia, AlphaSense). We do — methodology, cohort, and per-case results all public, reproducible from a fresh git clone in under 30 minutes. We also tested ourselves head-to-head against vanilla frontier models (GPT-5, Claude Sonnet 4.6, Gemini 3.1 Pro) on the same cohort and published the full table below. That's the discipline; the recall number is downstream of it.
Head-to-head against frontier vanilla
We ran the same 10 AAER cases through three vanilla frontier models (no taxonomy, no RAG, no Beneish — just a competent forensic-accountant prompt against the same md_a + risk_factors + audit_opinion text we send our own stack) and counted any pattern at severity medium or higher as a flag — exact same threshold our scorer uses.
| System | Recall | Flagged | 95% CI | Stack |
|---|---|---|---|---|
| OloLand v1.1 | 100% | 10/10 | [72%, 100%] | Beneish + EBITDA bridge + FORENSIC taxonomy (300+ checks) |
| GPT-5 (vanilla) | 90% | 9/10 | [60%, 98%] | Raw model + forensic prompt only |
| Claude Sonnet 4.6 (vanilla) | 80% | 8/10 | [49%, 94%] | Raw model + forensic prompt only |
| Gemini 3.1 Pro (vanilla) | 80% | 8/10 | [49%, 94%] | Raw model + forensic prompt only |
How to read this table. At n=10 the confidence intervals overlap heavily — don't read this as "OloLand beat GPT-5." Read it as the only published forensic-recall head-to- head in the M&A AI space, with the full per-case table below so you can check exactly which cases each system caught and missed.
What the table shows. Seven cases are universal flags (any competent forensic prompt catches them). One case (Becton Dickinson) was missed by all three vanilla backends and flagged by OloLand only at the very edge of the medium-severity threshold (1 pattern) — fragile, may flip on re-run. The architectural moat isn't the recall number on this small cohort; it's the platform integration that runs forensic extraction at scale on real deals with structured engines, state, audit trails, and the discipline to publish this table at all.
The architectural unlock between v1.0 (10%) and v1.1 (100%) was a single mode-flag in the production risk extractor (300+ leaf-level checks): RiskExtractionMode.FORENSIC swaps the active-DD "concrete evidence required" guardrail with a confidence-graded pattern-detection rubric. Production default behavior is byte-identical (verified) — only the benchmark and the Pre-LOI Forensic Screen capability opt in.
Methodology
Each case is an SEC AAER from 2020–2024 in OloLand's ICP (healthcare services, industrial distribution, B2B services, SaaS, specialty finance). For each, we identified the most recent 10-K filed before the first public fraud disclosure date, ran OloLand's v1.0 forensic battery on that filing's data only, and recorded whether any included engine fired.
Read the full methodology →Per-Case Results
| Company | Sector | AAER | Flagged | Triggered Engines |
|---|---|---|---|---|
| Gartner, Inc. | B2B Services | #4411 | ✓ | risk_taxonomy:9 |
| Cantaloupe, Inc. | SaaS | #4417 | ✓ | risk_taxonomy:27 |
| Co-Diagnostics, Inc. | Healthcare | #4428 | ✓ | beneish, risk_taxonomy:3 |
| GTT Communications, Inc. | B2B Services | #4459 | ✓ | risk_taxonomy:2 |
| Evoqua Water Technologies Corp. | B2B Services | #4493 | ✓ | risk_taxonomy:6 |
| HF Foods Group Inc. | Industrial | #4506 | ✓ | risk_taxonomy:8 |
| CPI Aerostructures, Inc. | Industrial | #4510 | ✓ | risk_taxonomy:22 |
| Moog Inc. | Industrial | #4532 | ✓ | risk_taxonomy:3 |
| Elanco Animal Health Inc. | Healthcare | #4537 | ✓ | risk_taxonomy:3 |
| Becton, Dickinson and Company | Healthcare | #4547 | ✓ | risk_taxonomy:1 |
Recall trend
Recall over the last 26 weeks. As the institutional-learning flywheel accumulates analyst corrections + promotes new model checkpoints, the line trends up.
Honest Limitations
- Selection: ICP-stratified, NOT a random sample. The current cohort is 10 AAERs across 5 ICP sectors (healthcare services, industrial distribution, B2B services, SaaS, specialty finance). Expansion to n=25-50 is on the roadmap.
- Sample size: n=10. CI half-width is wide for small cohorts (95% CI roughly ±22 pts at n=10).
- LLM variance: the risk extractor's per-case output is non-deterministic. The published number is single-seed; a 3-seed mean ± stdev (matching the Vals benchmark protocol) is a roadmap item, not the current headline. Individual cases occasionally flip flagged/unflagged between seeds — re-running may move the recall ±10pt.
- Data scope: Recall reflects what's achievable from 10-K-only data. The Pre-LOI Forensic Screen on a real deal has access to the full document corpus including GL exports, tax returns, management projections.
Reproduce / Submit a System
git clone https://github.com/ololand-ai/olo5.git cd olo5 uv run --project eval python -m eval.restatement_recall.runner
Want to submit a different system to the leaderboard? See CONTRIBUTING.md in the repo.