Vals AI Finance Agent — Methodology
How OloLand’s verifier stack is scored against the Vals AI Finance Agent v1.1 public split.
Benchmark
The Vals AI Finance Agent benchmark is a 537-question dataset co-created by Vals AI, Stanford researchers, and a Global Systemically Important Bank. Tasks reflect real entry-level financial-analyst workflows: SEC filing research, projections, multi-document reconciliation. Each question ships with a reference answer and a structured rubric of correctness and contradiction operators.
The dataset is split into three parts:
- Public Validation: 50 questions, open under CC-BY-4.0 on HuggingFace.
- Private Validation: 150 questions, available for license.
- Test: 337 questions, held out, submission-only via platform.vals.ai.
All numbers on the OloLand benchmark page are computed on the public split. We are licensing the private set next to scale the sample size.
Stages
Stage A — frontier-model baseline
The unmodified vals-ai/finance-agent Python harness, four standard tools (web_search via Tavily, edgar_search via sec-api.io, parse_html_page, retrieve_information), Claude Opus 4.7 routed via Google Vertex AI. Three independent seeds.
Stage C v2.1 — OloLand verifier-stack augmentation
Same harness, same model, same prompt scaffold. Added: nine stateless deterministic OloLand tools and a system-prompt nudge that instructs the agent to call them for any quantitative claim. Three independent seeds.
The nine OloLand tools:
| Tool | What it does |
|---|---|
compute_cagr | (end / start)^(1/n) − 1, year-aware |
compute_ratios | Margins, leverage, coverage, FCF margin from line items |
reconcile_metrics | Cross-document reconciliation with source hierarchy CPA > tax > management > AI |
compute_beneish_m_score | Forensic earnings-quality screen (Beneish 1999, 8 components) |
compute_guidance_beat_miss | Beat/miss vs guidance; mode='margin' computes midpoint as numerator_mid / denominator_mid (not avg of endpoint percentages) |
compute_implied_metric | Base × pct → labeled implied dollar value with explicit presentation hint |
compute_dcf | Five-step FCF DCF: rev → EBITDA → EBIT → NOPAT → FCF; Gordon-growth terminal; mid-year discounting |
summarize_comp_set | Percentile rank, median, mean of target vs peer multiples |
extract_metric_series | Multi-period series structuring with year-aware CAGR + YoY |
Two runtime patches (the only deviations from upstream)
Both patches sit outside the agent loop and the scoring path:
- Vertex routing patch (
scripts/run_via_vertex.py) replacesmodel_library.providers.anthropic.AnthropicModel.get_clientwith a version that returnsAsyncAnthropicVertexwhenGOOGLE_CLOUD_PROJECTis set. Also rewrites the model slug to the Vertex@defaultform and setscustom_endpoint='vertex://'to suppress Anthropic-only beta headers (files-api-2025-04-14,interleaved-thinking-2025-05-14) that Vertex rejects. - Stage C tool-registry patch (
scripts/stage_c/run_with_ololand_tools.py) replacesfinance_agent.get_agent.get_agentwith a verbatim copy whoseavailable_toolsdict includes the 9 OloLand tools. No upstream fork.
Judge
Claude Sonnet 4.6 on Vertex (claude-sonnet-4-6@default, global endpoint). Per-criterion scoring with the rubric structure shipped in the public dataset.
The judge prompt was tightened with explicit numerical-tolerance rules to prevent inter-run scoring variance on rounding boundaries:
- Percentages: ±0.10 percentage points → match
- Basis points: ±0.5 bps → match
- Dollar values: ±1.0% → match
- Multipliers / ratios: ±2.0% → match
- Counts / integers: exact
- Dates: same calendar date in any consistent format
- Strings: semantic equivalence
Multi-seed protocol
Each stage was run 3 times independently. Each seed:
- Fresh Vertex AI session (no shared client state)
- Same harness, same prompt, same model (
claude-opus-4-7@default,temperature=0.0,max_turns=50,parallelism=4) - Independent retrieval calls (Tavily / sec-api.io)
- Scored once with the tightened judge
Per-category mean and standard deviation reported. n=3 is the minimum for estimating variance; future iterations will expand to 5–10 seeds.
Per-tool attribution (Gemini 3.1 Pro, leave-one-out, 3-seed)
To answer the question “which of the 9 OloLand tools is doing the work?”, we ran a leave-one-out ablation on Gemini 3.1 Pro at temperature 1.0: nine Stage C runs, each with exactly one tool removed, three independent seeds per ablation. Mean accuracy delta vs the full Stage C reference (86.15% ± 0.18%) gives each tool’s individual contribution.
| Tool removed | 3-seed mean | ± SD | Δ vs full | Signal |
|---|---|---|---|---|
compute_cagr | 82.81% | 1.25% | +3.33pt | workhorse |
extract_metric_series | 85.10% | 0.73% | +1.04pt | borderline |
reconcile_metrics | 85.72% | 2.20% | +0.42pt | noise floor |
compute_beneish_m_score | 85.87% | 3.17% | +0.27pt | noise floor |
compute_dcf | 86.30% | 1.95% | −0.16pt | noise floor |
compute_guidance_beat_miss | 86.71% | 1.73% | −0.56pt | noise floor |
compute_implied_metric | 87.19% | 1.60% | −1.04pt | noise floor |
compute_ratios | 87.45% | 0.44% | −1.30pt | negative* |
summarize_comp_set | 87.95% | 3.82% | −1.80pt | noise floor (huge SD) |
What the table shows
compute_cagris the workhorse (+3.33pt ± 1.25%). Effect size dwarfs noise. Most Vals questions involve growth-rate calculation; Gemini’s in-token CAGR arithmetic is the single biggest variance source the deterministic tool eliminates.compute_ratioscontributes −1.30pt ± 0.44% — the only statistically detectablenegative contributor. Per-question diagnosis traces this to a single failure mode: when the tool is available, Gemini reaches for it as a shortcut on multi-period questions and reports only the computed ratio instead of the underlying time series the rubric requires (mirror of the v2.3-fixed Beat-or-Miss midpoint problem). Fix candidate is a prompt addendum requiring time-series reporting alongside any computed ratio.- The remaining seven tools land within ±1 SD of the full-stack baseline — statistically indistinguishable from zero mean-accuracy contribution at n=50. Their contribution is to variance reduction: ablation runs have SD 0.4–3.8% while the full stack has SD 0.18%. The tools collectively cover each other’s edge cases; pulling one out restores significant per-question variance to the agent’s reasoning chain.
*Negative contribution under investigation; fix candidates listed in the project memo. The 9 tools are kept as a set in the current benchmark configuration; ablation guides future iteration, not retroactive claim revision.
Cost per seed
| Component | Stage A | Stage C v2.1 |
|---|---|---|
| Vertex AI: agent (Opus 4.7) on 50 questions | ~$6 | ~$8 |
| Vertex AI: Sonnet 4.6 judge | ~$2 | ~$2 |
| Tavily web_search | free tier | free tier |
| sec-api.io EDGAR search | $55/mo plan | $55/mo plan |
| Total per seed | ~$8 | ~$10 |
Limitations
- Public split only. 50 of 537 questions. The full set almost certainly contains harder questions; the leaderboard top score is 64.4% on the full 537. Translate confidence accordingly.
- Sonnet 4.6 judge, not Vals’ production grader. Methodologically aligned, not bit-identical. Cross-validation against the platform’s scoring is a future-work item.
- n=3 seeds. Adequate at the headline level, thin for per-category claims (especially n=3 and n=4 categories). Per-category SD is illustrative, not a confidence interval.
- One category regression: Financial Modeling −2.8pt. Diagnosed as reading-comprehension variance on metric-definition ambiguity (one TSM question: MoM vs YoY interpretation), not an arithmetic-tool failure. Fix candidate: refine the prompt nudge with a literal-phrasing-preference clause. Detail in the white paper.
References
- Benchmark: vals.ai/benchmarks/finance_agent
- Vals methodology: vals.ai/methodology
- Harness (open source): github.com/vals-ai/finance-agent
- Public dataset: huggingface.co/datasets/vals-ai/finance_agent_benchmark
- Verifiability thesis: Andrej Karpathy, Sequoia AI Ascent 2026 keynote — “LLMs and reinforcement learning automate what you can verify.”