Restatement Recall — Methodology
How OloLand's forensic battery is scored against SEC AAERs.
Selection: 8-gate filter
Each AAER must pass all 8 gates before it enters the eval set:
- Industry inside OloLand's ICP: healthcare services, industrial distribution, B2B services, SaaS, specialty finance
- Mid-market scale: $50M–$1B revenue at restatement
- Pre-restatement 10-K parseable on EDGAR (text-based, not OCR-corrupt)
- Public company with CIK on SEC EDGAR
- AAER announced 2020–2024 (modern reporting standards)
- Restatement event distinct from announcement (data must predate disclosure)
- Material financial restatement — AAER cites a restatement of historical financial statements, not just an FCPA bribery, books-and-records, or auditor-independence violation. Forensic engines score financial accuracy; clean-books bribery cases are out of scope.
- Domestic Form 10-K filer — company files Form 10-K, not Form 20-F. Foreign private issuers have different filing structure (no MD&A in standard form); v1.0 schema only handles 10-Ks. v1.1 may extend.
Per gate 6, the "pre-AAER 10-K" used for scoring is the most recent 10-K filed before the date the fraud was first publicly disclosed (often via SIC investigation 8-K), NOT the date the AAER itself was announced. Otherwise OloLand would be scored against 10-Ks that already reflect the restatement — info already public, unfair test.
Engine applicability table
| Engine | Runnable on 10-K only? | Why |
|---|---|---|
| Beneish M-Score | ✅ | Needs balance sheet + income statement; 10-K has both |
| EBITDA bridge | ✅ (when non-GAAP recon present) | Needs reported EBITDA + adjustment table; ~70% of modern 10-Ks |
| Structured risk taxonomy (300+ leaf-level checks) | ❌ (v1.0) | Engine is async + deal-scoped; v1.0 eval venv intentionally lightweight. v1.1 will ship a sync wrapper. |
| Cross-doc reconciliation | ❌ | Needs multi-source data (CPA + tax + management); 10-K alone is single-source |
| Benford's Law | ❌ | Needs GL transaction amounts; not in 10-K |
| Lapping detection | ❌ | Needs AR aging by customer + cash receipts; not in 10-K |
Threshold definitions
- Beneish: M-Score > -1.78 → flag. (Canonical manipulation threshold from Beneish 1999.)
- EBITDA bridge:
bridge.adjustment_ratio > 0.10→ flag. Engine computes ratio = |total_adjustments| / |reported_ebitda|. Only runs when non-GAAP reconciliation table is present.
Wilson 95% confidence interval
We use the Wilson score interval (Wilson 1927) rather than the normal-approximation interval. At small n (n=25) and proportions near 0 or 1, the normal-approximation interval becomes unreliable; Wilson's interval is bounded to [0, 1] and well-calibrated across the full range. It's the standard choice recommended by most statisticians for n < 100.
Limitations
- v1.0 cohort is n=10 (not 25 as originally specced) — Wilson 95% CI half-width is ±28pp on a 0% point estimate. v1.1 expands to 25.
- ICP-stratified across 4 sectors (industrial distribution / healthcare services / b2b services / saas). The specialty_finance bucket is empty in v1.0 (Malvern Bancorp dropped — acquired before filing the FY2022 10-K).
- v1.0 publishes recall achievable from 10-K-only data — the same data ANY system has at pre-LOI time. The Pre-LOI Forensic Screen on a real deal has access to the full document corpus including GL exports, tax returns, management projections.
- Binary recall metric. v1.1 will add concern-matched recall (does the engine that fired match the AAER's actual concern category?).
- v1.0 runs 2 active engines (Beneish, EBITDA bridge). The structured risk taxonomy (300+ leaf-level checks) is excluded; v1.1 ships a sync wrapper service.
- human_reviewer: claude_auto — v1.0 manifests were curated by Claude Opus and auto-accepted without human review. The traditional benchmark workflow is human-review-gated; v1.0 trades that gate for faster iteration. v1.1 will re-curate with founder review per case.
- LLM extraction is the bottleneck of v1.0 recall. The curator's 10-K truncation (80,000 chars) cuts off most income statement / balance sheet tables before they finish; quarterly financials end up sparse, starving the Beneish engine of multi-period series. The 0% v1.0 recall reflects this extraction floor more than it reflects engine quality. v1.1 raises truncation to 200,000 chars and adds a second extraction pass focused on financial-statement tables.