Augments Anthropic Cowork · Depth behind every finance plugin

Anthropic ships breadth. OloLand ships depth.

Anthropic ships the workflow surface — five finance verticals, ten named agents, eleven read-only data connectors. OloLand ships what the verticals don’t: deterministic DCF/LBO/Monte Carlo, forensic QoE primitives, a structured risk taxonomy (300+ leaf-level checks), cross-deal institutional memory, and the compliance hooks Anthropic’s vertical plugins leave empty. Same Cowork session. Underwriting-grade computation underneath.

Why / What / How

Three layers. Three audiences. One stack.

OloLand is the depth layer for Anthropic-powered finance work — the verifier stack and persistent system of record that turns Claude’s generic agent runtime into IC-defensible private-equity and M&A underwriting.

Why
The artifact a PE partner asks for at the IC

An IC-ready memo.

Not search results. Not a chat transcript. The terminal artifact your partner names by hand at the meeting — with source-linked numbers, every assumption ledgered, and citations defensible in front of LPs.

What
The category we define

A verifiable intelligence layer.

Hebbia and AlphaSense are intelligence layers — better search over your data room. Verifiable intelligence is a different category: inputs reconciled against the source hierarchy (CPA > tax > management > AI), outputs cited to source pages, conclusions blocked from the memo when evidence is missing.

How
The mechanism Claude doesn’t ship

Deterministic engines and a persistent deal record.

Claude Cowork can read documents and call tools. It cannot run a real DCF against typed financial values, run the Beneish/Benford/EBITDA-bridge forensic battery, or remember what your associate flagged on deal #14 when you start deal #201. OloLand ships both — and they compose with the Claude agents you already use.

Converging Evidence

Three Teams. One Conclusion.

11%

Karpathy AutoResearch

700 autonomous experiments improved eval accuracy by 11%. No model was changed. The harness did all the work.

+13.7pp

LangChain Terminal Bench

Structured tool orchestration moved agentic coding scores from 52.8% to 66.5%. Same underlying models.

+13 pts

M&A Agentic Benchmark

OloLand scored 124/125 vs. Claude Pro at 114/125 on the same Sonnet 4.6 model. Infrastructure changes shipped in 24 hours.

Same Cowork surface. Different depth.

Anthropic Ships Breadth. OloLand Ships Depth.

We’re not a competitor to Anthropic’s finance plugins. We’re the deterministic, persistent, citation-enforced depth layer that runs underneath them. Both compose. The PE associate already running Anthropic’s private-equity plugin invokes ours when the IC pushes back on the number.

Anthropic’s 2026 finance lineup

The workflow surface

  • Five vertical plugins (private-equity, financial-analysis, investment-banking, equity-research, wealth-management)
  • Ten named finance agents (Pitch Agent, Valuation Reviewer, GL Reconciler, KYC Screener, Earnings Reviewer…)
  • Eleven read-only data connectors (Daloopa, Morningstar, S&P Kensho, FactSet, Moody’s, PitchBook, LSEG, Aiera, Chronograph, Egnyte, MT Newswires)
  • Microsoft 365 add-ins for Excel, PowerPoint, Word, Outlook
  • hooks/hooks.json shipped as [] across all five vertical plugins
OloLand’s depth layer

Underwriting-grade computation

  • Deterministic DCF, LBO, Monte Carlo, Real Options, Comps engines (strict unit enforcement)
  • Forensic QoE primitives: Beneish, Benford, EBITDA bridge, journal-entry, lapping detection, covenant cascade
  • Structured risk taxonomy (300+ leaf-level checks) with industry overlays + ensemble scoring
  • Cross-document reconciliation (CPA > tax > structured-feed > management > AI hierarchy)
  • Middle-office assumption controls + server-side IC approval gate + approval evidence snapshot (new, May 2026)
  • Cross-deal flywheel: similar deals, outcome calibration, analyst-correction retraining
  • MaskablePPO war-game RL engine (16-quarter competitive simulation)
  • Three OloLand plugins: ololand-dd, ololand-forensic-qoe, ololand-compliance-hooks
Ground Truth + AI

The Fusion Table

A language model has the right column. OloLand has both.

DomainGround Truth (Structured)Probabilistic (LLM)Why Both
FinancialsDCF engine computes EVLLM extracts assumptions from 10-KLLM cannot compute; engine cannot read
RiskStructured risk taxonomy (300+ checks) quantifiesLLM identifies risks from documentsTaxonomy without identification is empty
ComplianceOFAC/CFIUS/HSR gates screenLLM contextualizes regulatory exposureGates catch deterministic violations
ProvenanceCitation anchors (page/cell)LLM produces grounded analysisAnalysis without anchors is unverifiable
ValuationMonte Carlo simulates distributionsLLM parameterizes from DD findingsParameters without simulation are point estimates
DecisionPlaybook rule engine fires verdictsLLM synthesizes IC narrativeRules without narrative are unexplained
Reproducible Eval

M&A Agentic Benchmark: AI vs AI

Real M&A due diligence deal (SigmaTron International take-private, FY2025 10-K). Identical prompts, no coaching, no re-prompting.

SystemQ1 (/25)Q2 (/25)Q3 (/25)Q4 (/25)Q5 (/25)Total
ChatGPT Pro (GPT-4.1)171815----50/75*
Claude Pro (Sonnet 4.6)2123232423114/125
OloLand (Sonnet 4.6)2425252525124/125

*ChatGPT session expired before Q4-Q5. GPT-4.1 was the active model at time of evaluation. This is a reproducible eval — not a formal benchmark — and we invite independent replication.

Capability Comparison

General-Purpose AI vs. OloLand Harness

CapabilityGeneral-Purpose AIOloLand Harness
Frozen metricsNoneDCF/LBO/Monte Carlo deterministic engines
Persistent stateLost on session closeFull deal lifecycle in PostgreSQL
Cross-deal learningNoneOutcome database + calibration
AuditabilityNonePage/cell-level provenance chain
Risk quantificationEssays300+ leaf-level checks, dollar impacts, correlation
Analytical chainIndependent essaysQ3 risk -> Q4 WACC -> Q5 coverage ratios
SecurityConsumer-gradeSOC 2 Type II (in progress), TEE-ready
Compliance screeningNoneOFAC, CFIUS, HSR automated gates
Scenario simulationThree point estimatesMonte Carlo with correlated distributions
Strategic modelingStrategy essaysWar game with 10,000+ simulations
Infrastructure-Driven Improvement

13 Points in 24 Hours. Zero Model Upgrades.

Every improvement came from harness changes — the model was identical throughout.

v1

Initial eval (risk prefetch broken)

2026-03-12

111/125
v2

Risk data prefetch fix

2026-03-12

117/125
+6 pts
v3

Risk category expansion + correlation logic

2026-03-13

124/125
+7 pts

Anthropic captures the session. OloLand captures the institution.

Same Cowork surface. Same Claude harness. Different depth — and a persistent deal record that lives across the portfolio, not the conversation. See OloLand under a live IC pushback.

Optional analytics

Help us improve the acquisition experience.

With your permission, Google Analytics, Google Ads, Cloudflare, and PostHog measure page visits and conversion paths. PostHog autocapture and session recording stay off. We do not send form contents, uploaded documents, email addresses, or phone numbers in behavioral events. This choice does not enable personalized ads or enhanced-conversion user data. Read our Privacy Policy.