ConstraintCognitive automation2025-05-20
On the Finance Agent Benchmark — 537 expert-authored questions over recent SEC filings — the best model, OpenAI o3, reached 46.8% accuracy at $3.79 per query
Financial analystoccupation page →Event date / reported
2025-05-20
Evidence stage
ConstraintFailure, rollback, regulation or cost is suppressing adoption. Can lower an assessment or widen its uncertainty.
Tasks this bears on
Building and updating models
Three-statement models, valuation, budget roll-ups — and updating them when the actuals come in.
Being augmented✓ Evidence-backed
Where this applies
Nine task categories from information retrieval to financial modelling, authored with experts from banks, hedge funds and private equity; the agents were given Google Search and EDGAR access. Public-filing research only — not an analyst working inside their own firm's models and internal data.
What this means
The agents had the tools an analyst has — search and the filings database — and still got fewer than half the expert-written questions right, at nearly four dollars a query. That places the current line inside the job rather than around it: retrieval and summarising hold up, the multi-step work of building a number from a filing does not, and cost is a live constraint at volume.
What it does not yet show
A benchmark is not an employment outcome, and this one is public filings only — it says nothing about an analyst working inside their own firm models and internal data, where the context is richer and the errors are more expensive. Nor does it hold still: the same harness scored differently a year later, so treat 46.8% as a dated reading, not a ceiling.
What you can check
Take one model you maintain and one assumption inside it — the growth rate, say. Ask a current model to derive that number from the filings alone, then check every figure it cites back to the document. What you are measuring is not whether it sounds right but how many of its numbers survive being traced.
Does it change the assessment?
No. The impact index is never moved by a single event. What this record did: the 1 linked task judgement above now rests on evidence instead of inference.
Source
arXiv 2508.00828 (Bigeard, Nashold, Krishnan, Wu) · verified 2026-09-11 · Claude (CTO/COO) — arXiv paper page read in full 2026-09-11 · interpreted 2026-09-11 · Claude (CTO/COO) 2026-09-11