CapabilityCognitive automation2024-02-27
The strongest model reached 58% accuracy on data-based statistical and causal reasoning questions and struggled to combine causal knowledge with data, the QRData benchmark found
Statisticianoccupation page →Event date / reported
2024-02-27
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Statistical analysis and modelling
Estimating, modelling and testing, and judging what the data can and cannot support.
Being augmented✓ Evidence-backed
Where this applies
An academic benchmark of 411 questions with data sheets from textbooks, online learning materials and academic papers, testing statistical and causal reasoning. The strongest model, GPT-4, achieved 58 percent accuracy; models had difficulty with data analysis and causal reasoning and struggled to use causal knowledge and the provided data together. It tests 2024-era models on textbook-style questions, not statisticians' work.
What this means
General models still get many statistical and causal questions wrong when they have to work from data.
What it does not yet show
A 2024 benchmark of textbook-style questions; newer models may score higher.
What you can check
Open arXiv:2402.17644 and find "achieves an accuracy of 58%".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Liu et al. — Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data (QRData), arXiv 2402.17644 (submitted 27 Feb 2024) · verified 2026-10-01 · Claude (VOLO agent) · interpreted 2026-10-01 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.