False info, real problems: Can accountants tame AI hallucinations?

Despite the rise of stringent governance and rigorous control structures, the infamous tendency of artificial intelligence to occasionally generate false information, also known as "hallucinations," remains stubbornly untamed by even the world's largest accounting firms. 

Processing Content

For example, in late 2025 Deloitte used AI to write a report for the Australian government that was later found to contain nonexistent academic papers, fake source material and a fabricated quote attributed to a federal court judge; a similar incident with the Canadian government followed a few months later. This year Deloitte was joined by fellow Big Four firms EY, which issued a report on cybersecurity threats to loyalty programs that was later found to have inconsistent data and fake citations; KPMG, whose report on the benefits of AI had factual inaccuracies about several companies; and PwC, which released several reports containing factual errors, fake citations and poor sourcing.

But the issue goes well beyond accounting, as AI hallucinations increasingly are seen in legal filings (almost 2,000 recorded instances so far), scientific research papers, medical records and imaging — even police reports. Despite all of these institutions' reputations for strict standards and oversight, none have been able to fully prevent AI-generated false information from escaping into the public. And recent research suggests this problem is not going away anytime soon or, really, at all. Increasingly it is understood that hallucination risk is a fundamental property of generative AI and tools that rely upon it: no hallucination risk, no generative AI. 

Robot Doh
jackie_vfx - stock.adobe.com

What becomes clear then is this problem will be with us for some time, so high-profile incidents such as these — many of which originate within organizations with resources that dwarf those of the average accounting firm — underscore the importance of meticulous AI oversight for all practices, from international megafirms to sole proprietors working from home. 

However, one might question just how much more oversight is needed. After all, each of the Big Four firms is well known for its commitment to responsible governance and quality control, and they quickly developed their own frameworks for responsible and trustworthy AI once the technology hit the mainstream. And yet even they, with all their expertise, all their procedures and all their controls were still unable to prevent hallucinated information from escaping the confines of their firm. If it happened to them, it can happen to anyone. But how exactly did it happen to them in the first place?

Avani Desai, head of Top 50 Firm Schellman, which specializes in technology assurance, said that when a large international firm with tons of resources lets hallucinated information reach external audiences, what likely happened was not a failure of technology but of governance. But, she said, it's a very understandable one. 

"LLM's are incredibly convincing. They don't say, 'I'm making this up.' They present fabricated information with confidence. That's why verification matters. My guess is somewhere along the process, people assumed someone else had already validated the content. And guess what: That's not an AI problem. That's a control design problem," she said. 

Drew Armanino, the artificial intelligence solutions practice lead for Top 25 Firm Armanino, raised a similar point in that what likely happened was someone trusted an output they shouldn't have, which is not a difficult mistake to make as the outputs are designed to be as convincing as possible. It's not unimaginable that someone would miss the hallucination even after double- or triple-checking the results. 

"My sense is, when something like this happens, the root cause goes back to a state of overreliance," he said. "You trusted an output that came from an AI process or workflows implicitly, and that's really easy for us to do. That's quite literally how these tools are designed."

Donny Shimamoto, the head of tech-focused accounting consultancy IntrapriseTechKnowlogies, said it's much like cybersecurity in that while it is a technology-related issue, it ultimately comes down to humans. It's not just about having steps to follow but making sure people follow those steps. He speculated that somewhere along the control process, someone did not follow the steps they were supposed to. 

"AI is like cybersecurity: The weakest link is the end user," he said. "It would appear the staff preparing the analysis or report failed to double-check the information provided to them by AI. However, the failures are also indicative of a systematic quality management issue in those firms too, because subsequent reviews also didn't catch the hallucinations. Many CPAs think quality management only applies to audit, but in reality, it applies to all services provided by a CPA firm." 

Wenzel Reyes, senior director of methodology and audit solutions at risk management software company provider MindBridge, noted that it may not even have been a matter of someone not being trained or not knowing what to do, but of suffering the entirely human problem of being stressed out and tired. 

"When they read it, 'Oh, this sounds credible, and I trust my team that they checked every single citation, footnote, reference and whatnot of this report.' But the associate doesn't always have time. They're stressed out. They've got to get out the door and don't have time to double-check every single reference," he said. 

The reality of hallucinations

Regardless of how it happened, there was strong agreement that it is unrealistic to say it will never, ever happen again. Hallucination risk was acknowledged as a fact of life, something that comes with generative AI overall. Anyone who claims an AI system will never produce false information is only demonstrating they have no idea how AI works. 

"I wouldn't trust anyone who made that claim," said Desai. "I live in this world, and cybersecurity leaders don't promise they'll never experience an incident. They promise they'll reduce risk, detect issues quickly and respond effectively. AI should be viewed the same way. Hallucinations are an inherent characteristic of today's generative AI systems. The question isn't whether they can happen. The question is whether your governance catches them before they impact a customer or a client."

Others raised similar points: To say a firm's AI will never hallucinate is about as realistic as saying a firm will always catch said hallucinations before they escape the firm, which is about as realistic as saying the firm will always catch any problem in any work it does. 

"Just as we can't say we will catch all financial statement errors or fraud, we can never say we can catch all hallucinations," said Shimamoto.

In the absence of certainty, and with the acknowledgement that hallucination risk is a constant, preventing AI-generated false information from reaching external audiences is less a matter of absolute control and more a matter of management and mitigation. But even here firms might need to make peace with the limits of their oversight. It cannot be all-encompassing — people cannot inspect and verify every single output that ever comes from an AI model, nor should they be expected to. 

"It's not going to be very practical to be criticizing, double-checking everything because what was the point of embracing this technology if you're going to double-check everything? But there's going to be a mechanism where the firm has to be comfortable," said Reyes. "What are the parts of the process where you can critically challenge the output?"  

While it can be tempting to seek a technological solution for this mitigation and management, overall there was strong agreement that what it comes down to ultimately is the firm's system of quality control and people's ability to follow it. And, according to Armanino, they should still have formal processes and procedures in place for when that system inevitably fails. Any firm that does not think that will happen, he said, is far too comfortable. 

"I think if you're a firm and you're under the impression that you've done everything right and put in policies and processes and feel totally comfortable we won't have any issues, that's probably not the right posture," he said. "You should be prepared. You should have protocol around this. … You don't want to get caught on the back foot. You should have a well-thought-out response because you're never going to get 100% perfect outputs from AI, just like you're not going to get 100% perfect outputs from people."

That being said, there was also strong agreement that it was utterly unacceptable to let AI-generated false information escape the firm, especially if it will wind up in a client's hands. While hallucination risk is a constant, and while it is impossible to say the controls for that risk will never fail, there is no acceptable rate for hallucinations in work products. While the FDA may tolerate certain levels of contamination in food, there can be no tolerance of contamination in a firm's deliverables. 

"Your process should be designed so hallucinations never make it to the client," said Desai. "Internally, that's different. If someone is brainstorming marketing ideas or drafting an outline, hallucinations are a much lower risk because there's another layer before publication. And we have to remember, risk is contextual. That's another thing ISO 42001 emphasizes. You don't govern every AI use case the same way. You apply controls based on impact."

How exactly do firms manage hallucination risk? What sorts of processes and procedures do they employ to specifically head off AI-generated false information from escaping the confines of the firm? Our next article explores the practical reality of hallucination control.


For reprint and licensing requests for this article, click here.
Technology Practice management Artificial Intelligence Data governance Corporate ethics Automation
MORE FROM ACCOUNTING TODAY
Load More