CapabilityCognitive automation2025-08-29
The best model reached 96.50% accuracy on clean invoices, 92.71% on scanned invoices and 87.46% on scanned receipts without task-specific training, Fraunhofer researchers found
Data entry clerkoccupation page →Event date / reported
2025-08-29
Evidence stage
CapabilityA demo, benchmark or paper shows the task can be done. Updates what the technology can do — not what employers will do.
Tasks this bears on
Verifying and correcting records
Checking entered data against the source and fixing errors.
Being augmented✓ Evidence-backed
Where this applies
A benchmark of extracting fields from invoices and receipts with language models, zero-shot, on open datasets. The authors report that Gemini 2.5 Pro achieved the highest accuracy across all three datasets: 87.46% on scanned receipts, 96.50% on clean invoices and 92.71% on scanned invoices. It measures field accuracy on public datasets, not use in accounts-payable departments.
What this means
Invoice fields can now be read with high accuracy without custom training — high enough to take the keying, not high enough to skip review.
What it does not yet show
A benchmark on public datasets; not a measure of deployment or of jobs.
What you can check
Open the invoice-processing benchmark on arXiv (2509.04469) and find "96.50% on Clean Invoices".
Does it change the assessment?
No — and this stage does not move it either. A "Capability" record is real evidence, but it does not upgrade a task judgement on its own. The 1 linked judgement above stand where they were.
Source
Berghaus, Berger, Hillebrand, Cvejoski and Sifa (Fraunhofer IAIS) — Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing, arXiv:2509.04469 (submitted 29 Aug 2025) · verified 2026-09-30 · Claude (VOLO agent) · interpreted 2026-09-30 · Claude (VOLO agent)
Primary source — published by the party that did this, or the authority of record. No co-signature needed.