CapabilityCognitive automation2023-09-25
Stories generated by language models passed 3–10 times fewer creativity tests than stories by professional authors, an expert assessment found
Writer and authoroccupation page →Event date / reported
2023-09-25
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Writing fiction and books
Drafting novels, stories and non-fiction books.
Being augmented≈ Platform inference
Where this applies
A study in which 10 creative writers assessed 48 stories written either by professional authors or by language models, using a set of creativity tests. The authors report that model-generated stories pass 3–10 times fewer of the tests than stories written by professionals, and that none of the models used as assessors correlates positively with the experts. It used 2023 models; co-authors include researchers at an AI company.
What this means
Judged by writers, prompted models' stories fell far short of professional fiction — the craft gap was large in 2023.
What it does not yet show
2023 models and a small set of stories; newer models and fine-tuning change the picture.
What you can check
Open the arXiv paper "Art or Artifice?" (2309.14556) and find "3-10X less TTCW tests".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Chakrabarty, Laban, Agarwal, Muresan and Wu (Columbia University and Salesforce AI Research) — Art or Artifice? Large Language Models and the False Promise of Creativity, arXiv:2309.14556 (CHI 2024; submitted 25 Sep 2023) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.