CapabilityCognitive automation2026-05-18
On BacktestBench, the best of 23 language models reached 67.41% overall accuracy, and volatility and Sharpe ratio remained 'disaster zones' for all of them
Quantitative analystoccupation page →Event date / reported
2026-05-18
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Backtesting and quantitative code
Writing and maintaining the code for models and backtests, and computing performance and risk metrics correctly.
Being augmented✓ Evidence-backed
Where this applies
A benchmark in which models turn strategy descriptions into reproducible backtests, evaluated on 23 language models. The paper reports that the top model reached an overall accuracy of 67.41%, and that while simpler metrics such as win rate see higher accuracy, complex statistical indicators such as volatility and Sharpe ratio remain 'disaster zones' across all models. It is a fixed benchmark, not production work.
What this means
The numbers that decide whether a strategy is any good — volatility, Sharpe ratio — are exactly where models still go wrong. Checking them remains a quant's job.
What it does not yet show
A fixed benchmark; it does not measure real research work.
What you can check
Open arXiv 2605.17937 (BacktestBench) and find "67.41%" in the paper.
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Wang, Yang, Wu et al. (Beijing Normal University) — BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting, arXiv 2605.17937 (submitted 18 May 2026) · verified 2026-09-29 · Claude (VOLO agent) · interpreted 2026-09-29 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.
This record is cited in