Restatement Recall — Methodology

How OloLand's forensic battery is scored against SEC AAERs.

Selection: 8-gate filter

Each AAER must pass all 8 gates before it enters the eval set:

  1. Industry inside OloLand's ICP: healthcare services, industrial distribution, B2B services, SaaS, specialty finance
  2. Mid-market scale: $50M–$1B revenue at restatement
  3. Pre-restatement 10-K parseable on EDGAR (text-based, not OCR-corrupt)
  4. Public company with CIK on SEC EDGAR
  5. AAER announced 2020–2024 (modern reporting standards)
  6. Restatement event distinct from announcement (data must predate disclosure)
  7. Material financial restatement — AAER cites a restatement of historical financial statements, not just an FCPA bribery, books-and-records, or auditor-independence violation. Forensic engines score financial accuracy; clean-books bribery cases are out of scope.
  8. Domestic Form 10-K filer — company files Form 10-K, not Form 20-F. Foreign private issuers have different filing structure (no MD&A in standard form); v1.0 schema only handles 10-Ks. v1.1 may extend.

Per gate 6, the "pre-AAER 10-K" used for scoring is the most recent 10-K filed before the date the fraud was first publicly disclosed (often via SIC investigation 8-K), NOT the date the AAER itself was announced. Otherwise OloLand would be scored against 10-Ks that already reflect the restatement — info already public, unfair test.

Engine applicability table

EngineRunnable on 10-K only?Why
Beneish M-Score✅Needs balance sheet + income statement; 10-K has both
EBITDA bridge✅ (when non-GAAP recon present)Needs reported EBITDA + adjustment table; ~70% of modern 10-Ks
Structured risk taxonomy (300+ leaf-level checks)❌ (v1.0)Engine is async + deal-scoped; v1.0 eval venv intentionally lightweight. v1.1 will ship a sync wrapper.
Cross-doc reconciliation❌Needs multi-source data (CPA + tax + management); 10-K alone is single-source
Benford's Law❌Needs GL transaction amounts; not in 10-K
Lapping detection❌Needs AR aging by customer + cash receipts; not in 10-K

Threshold definitions

  • Beneish: M-Score > -1.78 → flag. (Canonical manipulation threshold from Beneish 1999.)
  • EBITDA bridge: bridge.adjustment_ratio > 0.10→ flag. Engine computes ratio = |total_adjustments| / |reported_ebitda|. Only runs when non-GAAP reconciliation table is present.

Wilson 95% confidence interval

We use the Wilson score interval (Wilson 1927) rather than the normal-approximation interval. At small n (n=25) and proportions near 0 or 1, the normal-approximation interval becomes unreliable; Wilson's interval is bounded to [0, 1] and well-calibrated across the full range. It's the standard choice recommended by most statisticians for n < 100.

Limitations

  • v1.0 cohort is n=10 (not 25 as originally specced) — Wilson 95% CI half-width is ±28pp on a 0% point estimate. v1.1 expands to 25.
  • ICP-stratified across 4 sectors (industrial distribution / healthcare services / b2b services / saas). The specialty_finance bucket is empty in v1.0 (Malvern Bancorp dropped — acquired before filing the FY2022 10-K).
  • v1.0 publishes recall achievable from 10-K-only data — the same data ANY system has at pre-LOI time. The Pre-LOI Forensic Screen on a real deal has access to the full document corpus including GL exports, tax returns, management projections.
  • Binary recall metric. v1.1 will add concern-matched recall (does the engine that fired match the AAER's actual concern category?).
  • v1.0 runs 2 active engines (Beneish, EBITDA bridge). The structured risk taxonomy (300+ leaf-level checks) is excluded; v1.1 ships a sync wrapper service.
  • human_reviewer: claude_auto — v1.0 manifests were curated by Claude Opus and auto-accepted without human review. The traditional benchmark workflow is human-review-gated; v1.0 trades that gate for faster iteration. v1.1 will re-curate with founder review per case.
  • LLM extraction is the bottleneck of v1.0 recall. The curator's 10-K truncation (80,000 chars) cuts off most income statement / balance sheet tables before they finish; quarterly financials end up sparse, starving the Beneish engine of multi-period series. The 0% v1.0 recall reflects this extraction floor more than it reflects engine quality. v1.1 raises truncation to 200,000 chars and adds a second extraction pass focused on financial-statement tables.

Optional analytics

Help us improve the acquisition experience.

With your permission, Google Analytics, Google Ads, Cloudflare, and PostHog measure page visits and conversion paths. PostHog autocapture and session recording stay off. We do not send form contents, uploaded documents, email addresses, or phone numbers in behavioral events. This choice does not enable personalized ads or enhanced-conversion user data. Read our Privacy Policy.