CapabilityCognitive automation2025-10-27
GPT-4o answered 83.69% of a national anaesthesiology exam correctly but did worse on application and analysis, and unsupported medical claims were its most common error
Anaesthesiologist / anaesthetistoccupation page →Event date / reported
2025-10-27
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Pre-operative assessment and planning
Examining the patient, reviewing history, medicines and risks, and deciding the anaesthetic plan.
Being augmented≈ Platform inference
Where this applies
A study running GPT-4o 30 times on Chile's 183-question anaesthesiology certification exam. Overall accuracy was 83.69%, highest on understanding and recall and lower on application (76.83%) and analysis (76.54%). Among incorrect answers, unsupported medical claims were the most common error. The authors say its limits in higher-order reasoning and diagnostic judgement call for more safeguards before clinical use.
What this means
A language model knows most of the exam, and is weakest where judgement is needed.
What it does not yet show
An exam, not patients; accuracy on questions is not safety in theatre.
What you can check
Open the BMC Medical Education study (doi 10.1186/s12909-025-08084-9) and find "overall accuracy of 83.69%".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Altermatt et al. — Evaluating GPT-4o in high-stakes medical assessments: performance and error analysis on a Chilean anesthesiology exam, BMC Medical Education 25 (27 October 2025) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.