Deloitte got caught. Not by a regulator. Not by internal quality review. By a researcher who actually read the report.
It's tempting to stop there and enjoy it. Don't. Deloitte isn't the story. The closed loop is — the same one I
Since then, investigations by the AI-detection firm GPTZero have caught
Deloitte said the report's conclusions held up even with the fabricated material removed. Maybe that's true. But "did the outcome change" is only the first of several questions a failure like this actually raises, and it's the only one Deloitte addressed publicly.
Was the evidence genuine? Does the process itself reveal a reliability problem that will resurface on the next engagement, regardless of outcome? Could an independent reviewer reconstruct the conclusion from the record alone, rather than being told to trust it? Deloitte never addressed any of the three it was actually asked. Neither has EY, KPMG or PwC. A firm that fabricates a citation once hasn't demonstrated it won't do it again — only that this time, someone caught it.
The profession has been here before, and it didn't fail on competence then either. Arthur Andersen's people were, by every measure, excellent — trained inside a culture built on telling clients no. What failed was institutional structure: Andersen's consulting revenue from Enron grew to exceed its audit fees from the same client, one firm answering to itself on both sides of the relationship. The Financial Accounting Standards Board has had eight chairs since 1973; seven came from the Big Eight, Big Six or Big Four, and the three most recent transitions run Deloitte to EY to KPMG, unbroken since 2013. None of that proves any standard was written to favor the industry it governs. It does mean the body writing the rules has never been chaired by anyone whose formation happened outside the firms those rules apply to.
The insurance industry has no stake in being right about any of this, and it has already priced the risk anyway. Standardized AI-exclusion language, developed by ISO and Verisk and adopted across most of the market, took effect industry-wide on Jan. 1, 2026. Hamilton Insurance Group writes specific products into the exclusion by name — ChatGPT, Bard, Midjourney, DALL-E. An underwriter at Aon put the question to Accounting Today directly: "Do you police it? Do you have protocols in place?" That's what a firm with no interest in being early or cautious is now asking accounting firms at renewal.
Scott Davis, a partner at Prager Metis, published the
He's right about how financial statement audits work. The comparison breaks down on the assumption that an AI system fails in the same predictable ways Excel does. It doesn't. Excel doesn't fabricate a court quote. It doesn't invent a statute and cite it with the same fluent confidence it would use for a real one. Deloitte's failure wasn't in the numbers — it was in the tool's account of what supported them, the kind of failure that testing the underlying transactions doesn't catch, because the falsehood was never in the transactions to begin with. A reader in South Africa reached the identical conclusion independently, with no stake in how this argument turns out. Independent convergence strengthens the inference.
This isn't written to score a point off Davis, Deloitte or any firm named above. Accounting Today's pages remain open the same way they were in June — I'd rather be corrected in public than right in a vacuum.
None of this requires auditing an entire AI system, any more than an audit today requires examining every transaction a client runs through its books. When you audit a bank processing millions of transactions daily, you don't examine each one — you examine whether the controls governing the system are reliable and independently verified.
The AI equivalent doesn't require opening a company's proprietary model to public inspection either. Escrowed, cryptographically attested, auditor-controlled testing can answer the specific question a Deloitte-style failure raises — did this citation actually exist — without exposing a single weight or line of training data. The auditor gets access proportional to what the specific work product requires: the citations, the numbers, the conclusions relied upon, not the whole system.
One requirement matters more than the mechanics of access: Whoever performs the verification must be competent to do it. A junior staffer confirming a citation exists is not the same as someone trained to know whether it says what the report claims. The system that produced a citation also can't be the one that verifies it — asking a model to grade its own homework tests nothing. Both conditions sound obvious. Neither is currently required anywhere in this profession's existing AI guidance.
The profession built that architecture once already, after the last time self-certification failed at scale, and it didn't wait for a second Enron to act. It won't need a formal AI Assurance Agency to start. The first step is smaller: Treat AI-generated authorities the way auditors already treat every other management representation — something to be verified, not presumed.
The Big Four aren't the villain in this. Deloitte, EY, KPMG and PwC are just the first ones the closed loop caught in public. Whether they're the last depends on whether a profession that's rebuilt this exact architecture once before is willing to do it again now, before the next documented failure arrives with a larger number attached than $290,000.









