AT Think

The most useful question about AI in audit

We all know the prevailing story: artificial intelligence is coming for the work, the drudgery and the judgment alike, and the only question is when. In audit, our experience building AI tools for the profession points the other way. As the tools take on more of the evidence-gathering, the human decision about what that evidence proves becomes the most valuable, and the most exposed, act in the engagement.

Processing Content

We don't see auditors becoming less integral to the work. We see them becoming more responsible for an audit's value.

The reason is structural. An audit produces assurance, not throughput. Investors, lenders and regulators rely on the opinion when they commit capital, and that opinion carries a named person's signature and a named person's liability. A tool that makes the work faster while making the conclusions harder to defend has not improved the audit. It has moved the risk to a place that is harder to see.

None of this is meant as a case against AI in audit. AI belongs there and already sits in the work for good reason. The argument is about how the tools are built, and about the one property that separates a tool that strengthens an audit from one that weakens it without anyone noticing.

The failure mode of audit is human, not machine

Start with the cautionary case. Wirecard, the German payments company, was added to the DAX in 2018 and filed for insolvency two years later, after a fraud built in large part on falsified documents and screenshots.

From 2016 to 2018, Wirecard's auditor verified the company's cash held abroad using documents and screenshots rather than contacting the banks directly. By the 2019 accounts, the balance in question had reached about €1.9 billion, said to sit in trustee accounts at two banks in the Philippines. Both the company and a third-party trustee supplied the confirmations. The auditor did not go to the banks itself. In 2020, when it finally did, the banks said the documents on their letterhead were false. Insolvency followed within days.

Every practitioner reading this knows the rule that was broken. Direct confirmation of bank balances is first-year training, and ISA 505 requires the auditor to control the external confirmation process. A screenshot supplied by the client carries no independent weight. It is management's own assertion offered as evidence for management's own assertion. The chain of evidence never left the party being audited, so it only ever confirmed itself.

Germany's audit oversight authority found the auditor had breached its professional duties, fined the firm €500,000 and barred it from taking on new public-interest clients for two years. Note what the sanction was for: a judgment failure, a decision to accept the client's own evidence in place of independent confirmation. The screenshots were only the mediums. Broken control was the cause.

Treating AI as the new risk in audit gets the history backwards. Bad evidence has always been a human failure, and automation changes only the speed and scale at which a weak control is applied. The Wirecard shortcut was efficient right up until it wasn't, because efficiency counts for nothing when the evidence isn't real. AI can make an audit faster. Whether it is true still depends on evidence that does not confirm itself.

More capable tools raise the judgment load

Better tools can make audits worse, not better, if they are built the wrong way. The research on automation bias is consistent: people over-weight confident automated output, verify it less, and inherit its errors, and the effect is strongest under time pressure and high stakes — which describes audit in busy season precisely. A commentary from the ICAEW put it plainly to practitioners: Depending on how a system is designed, an auditor can slide from active analysis of the data into passive review of an analysis the machine already made.

Anyone who has used a general-purpose AI assistant has felt the mechanism. You ask a question and get a fluent, confident answer that is sometimes simply wrong. If you catch it, it costs you a few minutes. If you don't, everything built on that answer is compromised. In a chat window, that is an annoyance. In an audit file, it can be fatal, and it is exactly the failure a reviewer is least likely to catch, because a confident wrong answer looks like a confident right one.

So, the question about an audit tool was never "how often is it right?" Average accuracy flatters the vendor and tells the auditor nothing about the single answer in front of them when they sign. What matters is whether the tool knows when it is right, whether it can tell a routine, high-confidence determination from one that needs a human, and whether it says so.

At issue is a design choice. A tool built for maximum autonomy answers everything with equal confidence and pushes the risk onto the reviewer. A tool built around calibration quantifies its own uncertainty and routes the low-confidence items to a person. It looks less impressive in a demo, because it visibly stops and asks. Stopping looks less impressive in a demo, because it visibly stops and asks. Stopping is what puts human attention where judgment is required, and it leaves a defensible record of who decided what.

At dnl, the impact is measurable. In the first months of 2026 our tool worked through 1.69 million disclosure-checklist questions across German, Austrian, IFRS and ESRS requirements. On the questions the system was confident enough to answer itself, reviewers accepted 98.17% without correction; human-prepared answers in the same workflow were accepted 92.68% of the time. The metric is reviewer acceptance, and the auto-answered questions are by design the ones the system judged it could stand behind — the harder ones went to people. That routing is the calibration doing its job, and any firm can test the same figures on its own engagements.

Different rules, but converging principles

Regulators are working on the same problem, and the direction is encouraging. No one has written an AI-specific audit rule, and the obligations that already bite come from the quality-management standards, which say nothing about artificial intelligence and apply to it completely. When the PCAOB opened its standard-setting agenda to public comment this summer, the large firms and vendors, ourselves included, largely urged staff guidance now rather than a multi-year standard that would be overtaken before it took effect. The U.K.'s FRC has issued AI-specific guidance, Germany's WPK maintains a working FAQ, and the IAASB has proposed revisions to its evidence standards. The paths differ; the principle does not.

The principle is that human accountability leads, and the technology serves it. Three properties make that concrete, and they are the questions a firm should ask of any audit tool it adopts. 

  • Traceability: Every AI-assisted conclusion can be traced to the specific evidence it rests on. 
  • Reproducibility: Inputs, model versions and outputs are documented well enough that a reviewer can re-perform the procedure, because re-performance is what inspection depends on. 
  • And human accountability: The record shows a human reviewed the work and remains responsible for the conclusion. The tool proposes. The auditor decides and signs.

Ultimately, the goal is a file — one an inspector can follow and a firm can defend. The choice in front of the profession is between an audit that is merely faster and one that is genuinely better, and only the second serves the investor who relies on it. The firms getting this right are building tools that put a sharper, better-supported human judgment at the center of the work. They judge the technology with a simple test: whether the resulting file is easier to defend. It's the same test  that made an audit trustworthy when the tools were bound books, ledgers and pencils.


For reprint and licensing requests for this article, click here.
Technology Audit Artificial Intelligence
MORE FROM ACCOUNTING TODAY
Load More