Third-party benchmarked · Two frontier models

OloLand on Vals AI Finance Agent v1.1

Same verifier stack, two frontier models from two different labs. OloLand collapses run-to-run variance 2.7× on Anthropic Claude Opus 4.7 and 11.8× on Google Gemini 3.1 Pro. Tools are doing the work, not the model.

Three independent seeds per condition, per model. Same harness, same 9 OloLand tools, same questions, same Sonnet-4.6 judge across both models — only the model swaps.

Public split (50 of 537 questions, CC-BY-4.0). Opus 4.7 Stage C v2.1 canonical 2026-05-09; Gemini 3.1 Pro Stage C v2.3 canonical 2026-05-10.

Variance reduction
2.7× / 11.8×
Opus 4.7 / Gemini 3.1 Pro
Mean accuracy lift
+3.4pt / +1.0pt
Opus 4.7 / Gemini 3.1 Pro
Per-tool workhorse
+3.3pt
compute_cagr alone (3-seed leave-one-out)

What the benchmark tests

The Vals AI Finance Agent benchmark is a 537-question dataset co-created by Vals AI, Stanford researchers, and a Global Systemically Important Bank. Tasks reflect real entry-level financial analyst workflows: SEC filing research, projections, multi-document reconciliation. Each question ships with a reference answer and a structured rubric of correctness and contradiction operators.

Anthropic’s Claude Opus 4.7 sits at #1 on the public leaderboard with 64.4%. We ran the leaderboard’s top model on the public 50-question split — three independent seeds per stage — with and without OloLand’s nine deterministic verifier-stack tools and a methodology-disambiguation prompt clause layered on top.

Vals AI Finance Agent v1.1 leaderboard →

Per-category lift

Mean accuracy ± standard deviation across 3 seeds. The largest lifts come precisely where deterministic computation replaces stochastic LLM arithmetic.

CategorynBaseline+ OloLandΔBase SDOloLand SD
Adjustments475.0%87.5%+12.5±0.0%±4.2%
Trends389.8%96.7%+6.9±9.1%±0.0%
Market Analysis352.4%58.5%+6.1±3.3%±4.4%
Complex Retrieval390.1%95.9%+5.8±5.4%±0.7%
Simple retrieval - Quantitative992.6%96.3%+3.7±6.4%±6.4%
Beat or Miss791.5%94.5%+3.1±4.1%±0.2%
Numerical Reasoning895.8%97.9%+2.1±7.2%±3.6%
Simple retrieval - Qualitative999.6%99.6%+0.0±0.4%±0.4%
Financial Modeling - Projections484.4%81.7%-2.8±2.2%±15.2%

The headline isn’t the ceiling. It’s the floor.

OloLand’s verifier stack doesn’t just lift accuracy — it collapses run-to-run variance. The Stage C mean is +3.3pt higher; the standard deviation is 2.7× narrower. The augmented agent isn’t just a higher expected-value bet, it’s a more predictable one.

For a private-equity IC memo or a quality-of-earnings screen, predictability is the load-bearing property — analysts need to know the answer they get this run is the answer they’d have got last run. With Trends, Opus 4.7 alone gets the right answer between 78.9% and 96.7% of the time depending on the seed. With OloLand’s tools, it’s 96.7% on every run — SD collapses from 9.1% to 0.0%.

Same story on Beat-or-Miss: SD 4.1% → 0.2%, a 20× tightening.

Methodology

  • Harness: The official vals-ai/finance-agent Python harness, unmodified except for a Vertex-AI routing patch and a tool-registry hook for the OloLand tools (no upstream fork).
  • Model: Claude Opus 4.7 routed via Google Vertex AI (claude-opus-4-7@default, temperature 0.0, max_turns 50, parallelism 4). Identical for both stages.
  • Tools (Stage C): 9 stateless deterministic tools (CAGR, ratios, cross-doc reconciliation, Beneish M-Score, beat/miss, implied metric, DCF, comp-set summary, metric-series structuring) + a methodology-disambiguation prompt clause.
  • Judge: Claude Sonnet 4.6 on Vertex AI, per-criterion scoring with the rubric structure shipped in the public dataset, and tightened numerical tolerances (percentages ±0.10pp, dollars ±1.0%, multipliers ±2.0%).
  • Multi-seed protocol: Three independent seeds per stage, fresh Vertex session each, scored by the same judge. Per- category mean and standard deviation reported.
Read the full methodology →
Verified on a second frontier model

Same lift on Gemini 3.1 Pro

We re-ran the entire protocol on Google’s Gemini 3.1 Pro — same harness, same OloLand tools, same questions, same judge, three independent seeds. Different model from a different lab. The verifier stack lifts both, and collapses run-to-run variance on each.

ModelStage A baseline+ OloLand verifier stackMean ΔVariance reduction
Anthropic Claude Opus 4.7
Vals leaderboard #1
89.4% ± 2.7%92.8% ± 1.0%+3.4pt2.7×
Google Gemini 3.1 Pro
Different lab, different family
85.1% ± 2.2%86.2% ± 0.18%+1.0pt11.8×

The cross-model invariant is variance collapse. Mean accuracy lifts differ by model (Opus +3.4pt, Gemini +1.0pt) — which is what you’d expect, since Opus has more headroom below its baseline. But on both models the verifier stack collapses run-to-run standard deviation by 2.7× (Opus) and 11.8× (Gemini).

On Gemini 3.1 Pro, three independent seeds of Stage C land at 86.17%, 86.32%, 85.95% mean accuracy — a 0.18-percentage-point spread. The same configuration on the same questions on the same model returns essentially the same number, every run. That’s what deterministic computation on top of a stochastic LLM looks like when measured.

Same harness, same prompts, same OloLand tools across both models. Only the model swaps. The tools are doing the work, not the model.

Honest limitations

  • Public split only. 50 of 537 questions. The full set (private validation 150 + held-out test 337) almost certainly contains harder questions; the leaderboard top score is 64.4% on the full 537. We are licensing the private validation set next.
  • n=3 seeds. Adequate for variance bands at the headline level, thin for per-category claims (especially in n=3 and n=4 categories). Per-category SD is illustrative, not a confidence interval.
  • Two models tested (Opus 4.7 and Gemini 3.1 Pro). Generalization to Sonnet 4.6, GPT-5.x, DeepSeek V4 is pending. Per-model prompt-nudge tuning was required: the Opus “MANDATORY use tools” framing over-applied on Gemini and was replaced with a softer “report all defensible representations” rule. Same 9 tools, same harness, same questions, same judge across both models.
  • Sonnet 4.6 judge, not Vals’ production grader. Methodologically aligned but not bit-identical. Cross-validation against the platform’s scoring is a future-work item once the private set is licensed.
  • One category regression: Financial Modeling −2.8pt. Diagnosed as reading-comprehension variance on metric-definition ambiguity, not a tool failure. Detail in the white paper.

Reproduce

Every figure on this page traces to a file in our eval repository. Both stages run end-to-end in about 15 minutes per seed for ~$8–10.

# Stage A baseline (one seed)
cd eval/vals_finance_agent
GOOGLE_CLOUD_PROJECT=super-agent-007 VERTEX_REGION=global \
  bash scripts/run_baseline.sh

# Stage C v2.1 (one seed)
GOOGLE_CLOUD_PROJECT=super-agent-007 VERTEX_REGION=global \
  bash scripts/stage_c/run_stage_c.sh

# Score with tightened judge
GOOGLE_CLOUD_PROJECT=super-agent-007 VERTEX_REGION=global \
  uv run --project ../../backend python scripts/score_local.py \
    --run-dir runs/<run_tag>_<TS>

Full white paper with per-seed numbers, judge prompt, tool implementations, and source-file pointers for every figure: docs/research/2026-05-08-vals-finance-agent-baseline.md

Stage C v2.1 canonical result — 2026-05-09. Anthropic ships the policy gradient. OloLand ships the grader.

Optional analytics

Help us improve the acquisition experience.

With your permission, Google Analytics, Google Ads, Cloudflare, and PostHog measure page visits and conversion paths. PostHog autocapture and session recording stay off. We do not send form contents, uploaded documents, email addresses, or phone numbers in behavioral events. This choice does not enable personalized ads or enhanced-conversion user data. Read our Privacy Policy.