CapabilityCognitive automation2024-04-18
Handing travel planning to a formal solver lifted success to 93.9%, against 10% for o1-preview alone, a study reported
Travel agent / advisoroccupation page →Event date / reported
2024-04-18
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Building complex itineraries
Fitting flights, hotels, transfers and budgets into one plan that meets every constraint.
Being augmented✓ Evidence-backed
Where this applies
A study that has a language model translate travel requests into constraint-satisfaction problems for formal solvers. On TravelPlanner it reports the best model, o1-preview, finding viable plans 10% of the time on its own, and the framework reaching a 93.9% success rate, with generalisation to unseen constraints. A test environment with a solver, not live booking.
What this means
Complex itineraries become solvable when the model hands the constraints to a solver — the hard part moves to stating the constraints correctly.
What it does not yet show
A research framework in a benchmark; it does not show deployment in agencies.
What you can check
Open arXiv 2404.11891 and find "success rate of 93.9%".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Hao et al. — Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools (arXiv 2404.11891, submitted 18 Apr 2024) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.