CapabilityCognitive automation2024-10-09
OpenAI built a benchmark of 75 Kaggle competitions to test agents at ML engineering; the best setup reached bronze-medal level in 16.9% of them
Machine learning engineeroccupation page →Event date / reported
2024-10-09
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Building a model for the problem
Picking an architecture, training it on your data, tuning it until the numbers move.
Automating≈ Platform inference
Deciding what counts as good enough
Building the test set, choosing the metric, and saying whether this thing may be used on real people yet.
Still human-led≈ Platform inference
Where this applies
Seventy-five ML-engineering Kaggle competitions - training models, preparing datasets, running experiments - graded against the human baselines on each competition's own public leaderboard, so the comparison is with the people who actually entered. The best-performing setup at the time of writing (o1-preview with the AIDE scaffold) reached at least bronze in 16.9% of competitions; the authors also report that resource scaling and pre-training contamination both change the number, and the benchmark is open-sourced so anyone can rerun it on a newer model. Two limits are structural rather than incidental: a Kaggle competition arrives with the problem already framed, the metric already chosen and the data already collected, which is the part of this job the benchmark cannot test at all; and OpenAI built the benchmark and sells models measured on it.
What this means
There is now a public, rerunnable measurement of the part of this job that looks most like a contest: a framed problem, a chosen metric, a prepared dataset, and a leaderboard of real entrants to lose to. On that part, agents were at bronze level in about one competition in six at the end of 2024, and the benchmark is open so the number moves in public rather than in a vendor's slide.
What it does not yet show
Bronze in a competition is not a model in production, and the benchmark cannot see the steps that decide whether one gets there: who framed the question, why this metric and not the other one, whether the training data is allowed to be used this way, and what happens on the day the distribution shifts. Nor does the number say anything about hiring. Read the date with it — it is a measurement of one model generation, and the authors themselves flag pre-training contamination as something that moves it.
What you can check
Take the last model you shipped and write down who chose the metric it is graded on. If that person was you, the part of your job this benchmark tests is not the part you are paid for. If it was somebody else, it is worth asking how they chose.
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 2 linked judgements above stand where they were.
Source
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering (OpenAI, arXiv:2410.07095) · verified 2026-09-12 · Claude (VOLO agent) — arXiv abstract page read; the 75-competition count, the 16.9% bronze figure and the AIDE/o1-preview pairing taken from the authors' own abstract · interpreted 2026-09-12 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.