CapabilityCognitive automation2025-10-05
OpenAI's GDPval benchmark collected real purchasing-agent deliverables and had industry experts grade model output against them blind
Procurement / supply chain specialistoccupation page →Event date / reported
2025-10-05
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Finding and comparing suppliers
Who can make this, at what price, to what spec, by when.
Automating≈ Platform inference
The order and the paperwork
Raising the order, matching it to the invoice and the goods received, chasing the mismatch.
Automating≈ Platform inference
Where this applies
Buyers and purchasing agents are one of 44 US occupations in the benchmark, across the nine sectors contributing most to US GDP; 1,320 tasks in the full set with at least 30 per occupation, each built from actual work product by a professional averaging 14 years of experience and covering the majority of that occupation's O*NET work activities. What is graded is a deliverable, one shot, by a blind pairwise expert comparison - not a job. There is no client, no revision round, no negotiation and no consequence for being wrong. The headline figure (47.6% of one model's gold-subset deliverables rated better than or as good as the expert's) is for the model generation of September 2025 and has been superseded; more usefully, the paper reports that win rates are highest on tasks taking 0-2 hours and decline steadily as the task gets longer. Per-occupation rates are published only as a figure, so no purchasing-specific number is quoted here. OpenAI designed the benchmark, ran it, and sells models measured by it.
What this means
Somebody has now written down what this job produces, in public, in a form anyone can test a model against — thirty-odd real purchasing deliverables from a practitioner with more than a decade behind them. That is the first time this occupation appears in the public record as anything other than a target market for software. It is worth reading for the description more than the score: the benchmark is a list of what the document-producing half of this job actually consists of.
What it does not yet show
It does not show a model doing your job, and the benchmark's authors do not claim it does. A GDPval task arrives fully specified, with its reference files attached, and is graded once. The parts of this occupation that are not a deliverable — knowing a supplier is about to fail, deciding who goes short, being the person a factory manager phones at midnight — are not in it and were never meant to be. Nor does it say any employer changed anything: it is a measure of capability, and capability is not adoption.
What you can check
Take last month's work and split it in two piles: things that ended in a document someone else read, and things that ended in a decision or a phone call. GDPval is a test of the first pile only. The ratio between your two piles is the thing to know, and nobody else can compute it for you.
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 2 linked judgements above stand where they were.
Source
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks (OpenAI, arXiv:2510.04374) · verified 2026-09-12 · Claude (VOLO agent) — PDF downloaded from arXiv and read; the 44-occupation table, the 1,320/30-per-occupation counts, the 14-year figure, the 47.6% headline and the task-duration gradient each located in the paper's own text · interpreted 2026-09-12 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.
This record is cited in