CapabilityCognitive automation2024-07-01
A world-class novelist outscored GPT-4 Turbo in literary critics' ratings of short stories, and the researchers concluded models are still far from challenging top writers
Writer and authoroccupation page →Event date / reported
2024-07-01
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Writing fiction and books
Drafting novels, stories and non-fiction books.
Being augmented≈ Platform inference
Where this applies
A contest between the novelist Patricio Pron and GPT-4 Turbo writing short stories from the same titles, rated by literature critics and scholars in English and Spanish. The authors report that language models are still far from challenging a top human creative writer, and that GPT-4 wrote more creatively when given the novelist's titles than its own. It compares one author and one model.
What this means
At the top of the craft, a novelist still clearly beat the model — and the model did better with the novelist's ideas than its own.
What it does not yet show
One novelist against one 2024 model; it says nothing about average writers or newer models.
What you can check
Open the arXiv paper "Pron vs Prompt" (2407.01119) and find "still far from challenging a top human creative writer".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Marco, Gonzalo, Mateo and del Castillo (UNED and Universidad Complutense de Madrid) — Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing?, arXiv:2407.01119 (EMNLP 2024; submitted 1 Jul 2024) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.