Restatement Recall

The public-algorithm floor for forensic AI on SEC restatements. Reproducible, peer-reviewable, openly methodology-disclosed.

v1.1 result: 100% of eligible SEC restatements caught from the pre-fraud- disclosure 10-K alone, using OloLand's full forensic stack (Beneish M-Score + EBITDA bridge + cross-doc reconciliation + structured risk taxonomy (300+ leaf-level checks) in RiskExtractionMode.FORENSIC).

n=10, 95% CI: 72%–100%, single-seed

Why publish this number?

Big-4 QoE shops don't publish recall numbers. Neither do AI competitors (Hebbia, AlphaSense). We do — methodology, cohort, and per-case results all public, reproducible from a fresh git clone in under 30 minutes. We also tested ourselves head-to-head against vanilla frontier models (GPT-5, Claude Sonnet 4.6, Gemini 3.1 Pro) on the same cohort and published the full table below. That's the discipline; the recall number is downstream of it.

Head-to-head against frontier vanilla

We ran the same 10 AAER cases through three vanilla frontier models (no taxonomy, no RAG, no Beneish — just a competent forensic-accountant prompt against the same md_a + risk_factors + audit_opinion text we send our own stack) and counted any pattern at severity medium or higher as a flag — exact same threshold our scorer uses.

SystemRecallFlagged95% CIStack
OloLand v1.1100%10/10[72%, 100%]Beneish + EBITDA bridge + FORENSIC taxonomy (300+ checks)
GPT-5 (vanilla)90%9/10[60%, 98%]Raw model + forensic prompt only
Claude Sonnet 4.6 (vanilla)80%8/10[49%, 94%]Raw model + forensic prompt only
Gemini 3.1 Pro (vanilla)80%8/10[49%, 94%]Raw model + forensic prompt only

How to read this table. At n=10 the confidence intervals overlap heavily — don't read this as "OloLand beat GPT-5." Read it as the only published forensic-recall head-to- head in the M&A AI space, with the full per-case table below so you can check exactly which cases each system caught and missed.

What the table shows. Seven cases are universal flags (any competent forensic prompt catches them). One case (Becton Dickinson) was missed by all three vanilla backends and flagged by OloLand only at the very edge of the medium-severity threshold (1 pattern) — fragile, may flip on re-run. The architectural moat isn't the recall number on this small cohort; it's the platform integration that runs forensic extraction at scale on real deals with structured engines, state, audit trails, and the discipline to publish this table at all.

The architectural unlock between v1.0 (10%) and v1.1 (100%) was a single mode-flag in the production risk extractor (300+ leaf-level checks): RiskExtractionMode.FORENSIC swaps the active-DD "concrete evidence required" guardrail with a confidence-graded pattern-detection rubric. Production default behavior is byte-identical (verified) — only the benchmark and the Pre-LOI Forensic Screen capability opt in.

Methodology

Each case is an SEC AAER from 2020–2024 in OloLand's ICP (healthcare services, industrial distribution, B2B services, SaaS, specialty finance). For each, we identified the most recent 10-K filed before the first public fraud disclosure date, ran OloLand's v1.0 forensic battery on that filing's data only, and recorded whether any included engine fired.

Read the full methodology →

Per-Case Results

CompanySectorAAERFlaggedTriggered Engines
Gartner, Inc.B2B Services#4411✓risk_taxonomy:9
Cantaloupe, Inc.SaaS#4417✓risk_taxonomy:27
Co-Diagnostics, Inc.Healthcare#4428✓beneish, risk_taxonomy:3
GTT Communications, Inc.B2B Services#4459✓risk_taxonomy:2
Evoqua Water Technologies Corp.B2B Services#4493✓risk_taxonomy:6
HF Foods Group Inc.Industrial#4506✓risk_taxonomy:8
CPI Aerostructures, Inc.Industrial#4510✓risk_taxonomy:22
Moog Inc.Industrial#4532✓risk_taxonomy:3
Elanco Animal Health Inc.Healthcare#4537✓risk_taxonomy:3
Becton, Dickinson and CompanyHealthcare#4547✓risk_taxonomy:1

Recall trend

Recall over the last 26 weeks. As the institutional-learning flywheel accumulates analyst corrections + promotes new model checkpoints, the line trends up.

No benchmark runs recorded yet. The recall trend will appear here once the weekly drift check has run a few times.

Honest Limitations

  • Selection: ICP-stratified, NOT a random sample. The current cohort is 10 AAERs across 5 ICP sectors (healthcare services, industrial distribution, B2B services, SaaS, specialty finance). Expansion to n=25-50 is on the roadmap.
  • Sample size: n=10. CI half-width is wide for small cohorts (95% CI roughly ±22 pts at n=10).
  • LLM variance: the risk extractor's per-case output is non-deterministic. The published number is single-seed; a 3-seed mean ± stdev (matching the Vals benchmark protocol) is a roadmap item, not the current headline. Individual cases occasionally flip flagged/unflagged between seeds — re-running may move the recall ±10pt.
  • Data scope: Recall reflects what's achievable from 10-K-only data. The Pre-LOI Forensic Screen on a real deal has access to the full document corpus including GL exports, tax returns, management projections.

Reproduce / Submit a System

git clone https://github.com/ololand-ai/olo5.git
cd olo5
uv run --project eval python -m eval.restatement_recall.runner

Want to submit a different system to the leaderboard? See CONTRIBUTING.md in the repo.

Restatement Recall v1.1 — 5/12/2026

Optional analytics

Help us improve the acquisition experience.

With your permission, Google Analytics, Google Ads, Cloudflare, and PostHog measure page visits and conversion paths. PostHog autocapture and session recording stay off. We do not send form contents, uploaded documents, email addresses, or phone numbers in behavioral events. This choice does not enable personalized ads or enhanced-conversion user data. Read our Privacy Policy.