CapabilityCognitive automation2024-02-02
On a benchmark of real-world travel planning, GPT-4 met all the constraints only 0.6% of the time, researchers found
Travel agent / advisoroccupation page →Event date / reported
2024-02-02
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Building complex itineraries
Fitting flights, hotels, transfers and budgets into one plan that meets every constraint.
Being augmented✓ Evidence-backed
Where this applies
A benchmark in which language agents must plan trips under many constraints using tools. The authors report that language agents are not yet capable of such complex planning tasks — even GPT-4 achieved a success rate of 0.6% — and that agents struggle to stay on task, use the right tools or keep track of multiple constraints. A test environment with 2023–24 models.
What this means
Holding a whole trip's constraints at once is exactly what language models were worst at.
What it does not yet show
A benchmark with 2023–24 models; newer systems do better, as later work shows.
What you can check
Open arXiv 2402.01622 and find "success rate of 0.6%".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Xie et al. — TravelPlanner: A Benchmark for Real-World Planning with Language Agents (arXiv 2402.01622, submitted 2 Feb 2024) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.