Vals AI Finance Agent — Methodology

How OloLand’s verifier stack is scored against the Vals AI Finance Agent v1.1 public split.

Benchmark

The Vals AI Finance Agent benchmark is a 537-question dataset co-created by Vals AI, Stanford researchers, and a Global Systemically Important Bank. Tasks reflect real entry-level financial-analyst workflows: SEC filing research, projections, multi-document reconciliation. Each question ships with a reference answer and a structured rubric of correctness and contradiction operators.

The dataset is split into three parts:

  • Public Validation: 50 questions, open under CC-BY-4.0 on HuggingFace.
  • Private Validation: 150 questions, available for license.
  • Test: 337 questions, held out, submission-only via platform.vals.ai.

All numbers on the OloLand benchmark page are computed on the public split. We are licensing the private set next to scale the sample size.

Stages

Stage A — frontier-model baseline

The unmodified vals-ai/finance-agent Python harness, four standard tools (web_search via Tavily, edgar_search via sec-api.io, parse_html_page, retrieve_information), Claude Opus 4.7 routed via Google Vertex AI. Three independent seeds.

Stage C v2.1 — OloLand verifier-stack augmentation

Same harness, same model, same prompt scaffold. Added: nine stateless deterministic OloLand tools and a system-prompt nudge that instructs the agent to call them for any quantitative claim. Three independent seeds.

The nine OloLand tools:

ToolWhat it does
compute_cagr(end / start)^(1/n) − 1, year-aware
compute_ratiosMargins, leverage, coverage, FCF margin from line items
reconcile_metricsCross-document reconciliation with source hierarchy CPA > tax > management > AI
compute_beneish_m_scoreForensic earnings-quality screen (Beneish 1999, 8 components)
compute_guidance_beat_missBeat/miss vs guidance; mode='margin' computes midpoint as numerator_mid / denominator_mid (not avg of endpoint percentages)
compute_implied_metricBase × pct → labeled implied dollar value with explicit presentation hint
compute_dcfFive-step FCF DCF: rev → EBITDA → EBIT → NOPAT → FCF; Gordon-growth terminal; mid-year discounting
summarize_comp_setPercentile rank, median, mean of target vs peer multiples
extract_metric_seriesMulti-period series structuring with year-aware CAGR + YoY

Two runtime patches (the only deviations from upstream)

Both patches sit outside the agent loop and the scoring path:

  1. Vertex routing patch (scripts/run_via_vertex.py) replaces model_library.providers.anthropic.AnthropicModel.get_client with a version that returns AsyncAnthropicVertex when GOOGLE_CLOUD_PROJECT is set. Also rewrites the model slug to the Vertex @default form and sets custom_endpoint='vertex://' to suppress Anthropic-only beta headers (files-api-2025-04-14, interleaved-thinking-2025-05-14) that Vertex rejects.
  2. Stage C tool-registry patch (scripts/stage_c/run_with_ololand_tools.py) replaces finance_agent.get_agent.get_agent with a verbatim copy whose available_tools dict includes the 9 OloLand tools. No upstream fork.

Judge

Claude Sonnet 4.6 on Vertex (claude-sonnet-4-6@default, global endpoint). Per-criterion scoring with the rubric structure shipped in the public dataset.

The judge prompt was tightened with explicit numerical-tolerance rules to prevent inter-run scoring variance on rounding boundaries:

  • Percentages: ±0.10 percentage points → match
  • Basis points: ±0.5 bps → match
  • Dollar values: ±1.0% → match
  • Multipliers / ratios: ±2.0% → match
  • Counts / integers: exact
  • Dates: same calendar date in any consistent format
  • Strings: semantic equivalence

Multi-seed protocol

Each stage was run 3 times independently. Each seed:

  • Fresh Vertex AI session (no shared client state)
  • Same harness, same prompt, same model (claude-opus-4-7@default,temperature=0.0, max_turns=50, parallelism=4)
  • Independent retrieval calls (Tavily / sec-api.io)
  • Scored once with the tightened judge

Per-category mean and standard deviation reported. n=3 is the minimum for estimating variance; future iterations will expand to 5–10 seeds.

Per-tool attribution (Gemini 3.1 Pro, leave-one-out, 3-seed)

To answer the question “which of the 9 OloLand tools is doing the work?”, we ran a leave-one-out ablation on Gemini 3.1 Pro at temperature 1.0: nine Stage C runs, each with exactly one tool removed, three independent seeds per ablation. Mean accuracy delta vs the full Stage C reference (86.15% ± 0.18%) gives each tool’s individual contribution.

Tool removed3-seed mean± SDΔ vs fullSignal
compute_cagr82.81%1.25%+3.33ptworkhorse
extract_metric_series85.10%0.73%+1.04ptborderline
reconcile_metrics85.72%2.20%+0.42ptnoise floor
compute_beneish_m_score85.87%3.17%+0.27ptnoise floor
compute_dcf86.30%1.95%−0.16ptnoise floor
compute_guidance_beat_miss86.71%1.73%−0.56ptnoise floor
compute_implied_metric87.19%1.60%−1.04ptnoise floor
compute_ratios87.45%0.44%−1.30ptnegative*
summarize_comp_set87.95%3.82%−1.80ptnoise floor (huge SD)

What the table shows

  • compute_cagr is the workhorse (+3.33pt ± 1.25%). Effect size dwarfs noise. Most Vals questions involve growth-rate calculation; Gemini’s in-token CAGR arithmetic is the single biggest variance source the deterministic tool eliminates.
  • compute_ratios contributes −1.30pt ± 0.44% — the only statistically detectablenegative contributor. Per-question diagnosis traces this to a single failure mode: when the tool is available, Gemini reaches for it as a shortcut on multi-period questions and reports only the computed ratio instead of the underlying time series the rubric requires (mirror of the v2.3-fixed Beat-or-Miss midpoint problem). Fix candidate is a prompt addendum requiring time-series reporting alongside any computed ratio.
  • The remaining seven tools land within ±1 SD of the full-stack baseline — statistically indistinguishable from zero mean-accuracy contribution at n=50. Their contribution is to variance reduction: ablation runs have SD 0.4–3.8% while the full stack has SD 0.18%. The tools collectively cover each other’s edge cases; pulling one out restores significant per-question variance to the agent’s reasoning chain.

*Negative contribution under investigation; fix candidates listed in the project memo. The 9 tools are kept as a set in the current benchmark configuration; ablation guides future iteration, not retroactive claim revision.

Cost per seed

ComponentStage AStage C v2.1
Vertex AI: agent (Opus 4.7) on 50 questions~$6~$8
Vertex AI: Sonnet 4.6 judge~$2~$2
Tavily web_searchfree tierfree tier
sec-api.io EDGAR search$55/mo plan$55/mo plan
Total per seed~$8~$10

Limitations

  • Public split only. 50 of 537 questions. The full set almost certainly contains harder questions; the leaderboard top score is 64.4% on the full 537. Translate confidence accordingly.
  • Sonnet 4.6 judge, not Vals’ production grader. Methodologically aligned, not bit-identical. Cross-validation against the platform’s scoring is a future-work item.
  • n=3 seeds. Adequate at the headline level, thin for per-category claims (especially n=3 and n=4 categories). Per-category SD is illustrative, not a confidence interval.
  • One category regression: Financial Modeling −2.8pt. Diagnosed as reading-comprehension variance on metric-definition ambiguity (one TSM question: MoM vs YoY interpretation), not an arithmetic-tool failure. Fix candidate: refine the prompt nudge with a literal-phrasing-preference clause. Detail in the white paper.

References

← Back to results

Optional analytics

Help us improve the acquisition experience.

With your permission, Google Analytics, Google Ads, Cloudflare, and PostHog measure page visits and conversion paths. PostHog autocapture and session recording stay off. We do not send form contents, uploaded documents, email addresses, or phone numbers in behavioral events. This choice does not enable personalized ads or enhanced-conversion user data. Read our Privacy Policy.