Auditor — tasks, one by one
The unit of analysis is the task, not the job title. Each one below carries its direction, whether the judgement rests on evidence or on platform inference, the reasoning, and what it does not establish.
Every task on this page#
Sampling and testing
Automating✓ Evidence-backedPulling a sample of transactions and checking them against the supporting documents.
Sampling exists because testing everything was too expensive. Once testing everything becomes cheap, the sample loses its reason to exist — and full-population testing has been ordinary in large audits for a decade, well before anything called AI. What generative tools add is the ability to read the supporting document rather than only the ledger entry, which extends the same trend into the half that used to require eyes.
Says nothing about how audit hours are actually billed, and in many firms the hours freed here were absorbed into more testing rather than into fewer people. It also says nothing about smaller audits, where the tooling is often not bought at all.
Deciding where to look
Still human-led≈ Platform inferenceRisk assessment: which accounts, which estimates, which part of this business could hide something.
Anomaly detection genuinely helps and finds things a person would miss. What it cannot do is know which unusual thing is worth pursuing in this business, this year, given this management's incentives — that requires a model of why someone would want to misstate, and the incentive lives outside the data.
A judgement about the nature of the work, not a measurement. It also does not claim auditors are good at this: missing the risk that mattered is the defining failure of this profession and there is a long public record of it.
Asking the question that gets a real answer
Still human-led≈ Platform inferenceSitting across from a finance director and working out what is not being said.
Evidence in an audit is not only documentary; a great deal of it is what someone does or does not say when asked directly. That is a live social act with a person who may have a reason to mislead, and the auditor's professional scepticism is exercised in the room, not in the file.
Nothing here measures how much of a modern audit is inquiry versus document testing, and in many engagements the inquiry is a formality completed by email. Where it is a formality, this task is not human-led — it is simply not being done.
The file and the opinion
Being augmented✓ Evidence-backedDocumenting what was done, and drafting the wording that goes in front of shareholders.
Working papers are structured, repetitive and heavily templated — close to the ideal case for generation, and this is where firms claim most of their time saving. The opinion itself is different in kind: its wording is regulated, and a modified opinion is a public act with consequences for the client that the person signing has to be willing to cause.
Nothing public measures documentation time saved in audit, and firm claims about it come from a party selling the transformation. What would settle it: a regulator's inspection findings on files prepared with these tools.
Signing it
Still human-led≈ Platform inferenceBeing the named person whose licence stands behind the opinion, and who answers if it was wrong.
Not a difficulty claim. In every market that has a statutory audit, the law requires a licensed individual or firm to sign, and attaches liability to that signature. That is the whole product: an audit is valuable precisely because someone can be sued over it. Automation does not reach a position that exists in order to be liable.
It protects the signature, not the headcount behind it. A firm can sign the same number of opinions with fewer people, and the signature says nothing about how many juniors were needed to get there — which is the number most readers of this page actually care about.
Owning what the machine drafted
New task✓ Evidence-backedChecking machine-drafted professional output before it leaves under the firm's name — citations, quotations, the things that look right.
New work created by the tooling, and the failure mode is specific: generated professional text is most dangerous where it is most confident, because a fabricated citation looks exactly like a real one to a reviewer who is skimming. This site holds a record of a Big Four firm repaying part of a government fee after fabricated references and a misattributed court quotation were found in a report it had drafted with a generative tool.
That record is one consulting engagement at one firm, not a statutory audit, and it establishes that the failure happened rather than how often it does. New work appearing is also not new headcount: in a billable-hours business an unbilled checking duty is absorbed, which is exactly the condition under which it gets skipped.