In our first article, we explored how AI-generated false information, also known as hallucinations, have increasingly been making their way into firm work products. Here we explore how some firms control specifically for this risk.
Accountants tend to view hallucination control as a governance challenge first and a technical challenge second. Essential to this governance is human review, which industry experts cite as key to preventing AI-generated false information from escaping the firm. And generally, the higher the consequences of a possible hallucination, the more human review is required.
"Anything involving technical conclusions, regulations, citations, legal interpretations, client recommendations or audit conclusions has to be validated by a human," said Avani Desai, CEO of Top 50 Firm Schellman, a tech assurance specialist. "AI can absolutely accelerate drafting, summarization, brainstorming or organizing information, but it doesn't become the source of truth."
Drew Armanino, artificial intelligence solutions practice lead for Top 25 Firm Armanino, described a similar concept. Workflows are designed around risk reduction through human review, with each step designed as a sort of mini-use case for AI that has its own risk management procedures. He raised the example of using an internal AI tool that helps develop technical memos. When breaking down the workflow for using this tool, he said the firm applied the same kind of logic one might apply to doing it manually. In such a case, the senior might set up the framework for what the facts are; then the staff accountants create the memo using prior examples or whatever other materials come in; then the manager does the pre-review; and then the partner does the final review. The firm applies much of that same logic into the actual AI workflow.

"There's literally a piece of the workflow that says, 'Senior, have you reviewed this?' and you have to check a little box that says, 'Yes, I have reviewed and I agree to the conclusions that are there.' Then you go to the next step," Armanino explained. "We have the same checks built into every step, so we're forcing that human interaction within the actual workflow. We're mirroring what would happen manually, and we're very intentionally forcing that human in the loop interaction."
Denny Ard, national director of innovation at Top 10 Firm Forvis Mazars, said multiple layers of control at his firm specifically target hallucination risk. He described an approach that emphasizes communication, planning and, like others, massive levels of human review at every step.
Policies require that all AI-generated content be reviewed by a human before it is used as support for engagements at Forvis Mazars. Specifically, outputs must be compared against source materials and applicable professional standards, then evaluated for completeness, accuracy and relevance. In addition, engagement teams must discuss planned AI use and professional skepticism during the planning process on the front end of a project.
On certain engagements, these discussions also include review of checklists where reviewers are required to specifically consider that AI outputs were reviewed and that any AI-identified discrepancies were addressed. Finally, all engagement deliverables — AI or not — are subject to an independent review process, where the final results are evaluated for reasonableness and compliance with applicable professional standards. Ard added that the firm's acceptable use policy requires disclosure when AI materially contributed to externally facing content, which mandates review of that output against Forvis Mazars' standards for accuracy and appropriateness before release.
"Our philosophy is straightforward: AI is a productivity and research aid that supports our professionals," said Ard. "It's a tool. It's never used as a substitute for professional judgment. We have built a framework of AI policies and guidance, and all of it is centered around one nonnegotiable principle: A qualified professional must review, validate and take ownership of anything that AI helps produce before it is relied upon."
But while a human in the loop is necessary for safe and trustworthy AI, it is not sufficient alone. Wenzel Reyes, senior director of methodology and audit solutions at risk management software company MindBridge, said organizations need to ask themselves specific questions about which humans they intend to have in the loop, and where they intend to place them.
"'Human in the loop' is such a complicated topic," he said. "Where do you put the human? The beginning of the process? The end of the process? In the middle? What skills should the human have? Is the human independent of the project? Because if I was the part of the project where I issued the report of course, I'm going to say I'm comfortable with the rest. I'm just going to check it at the very end."
Donny Shimamoto, head of tech-focused accounting consultancy IntrapriseTechKnowlogies, raised another question: Is the human reviewer even qualified to do the review? If a human cannot understand the work in the first place, how can they be expected to assess it?
"The failure to use a qualified reviewer is also a failure to meet our ethical standards. That would be like asking a tax partner to review an audit report. Both the firm and the reviewer need to consider whether they can properly assess the quality of the work product," he said.
Without asking these sorts of questions, it can be easy to slip into a sort of accountability theater, where people go through the motions of review but are mainly just rubberstamping whatever falls into their inbox.
Ard said Forvis Mazars is well aware of this risk and explicitly accounts for it in its policies. The firm's review process specifically rejects someone just looking over a work product and deciding it sounds reasonable before passing it on. Reviewers must confirm the content accurately reflects the underlying data or documents; that summaries are faithful to the original; that data references and citations are correct and current; and that no material facts or exceptions were omitted. Then, they must check the output against the broader engagement file and known client-specific risks. Finally, they must apply professional judgment by skeptically assessing whether conclusions are appropriate, whether any data transformations or formulas were applied correctly, and whether assumptions or limitations are clearly stated and addressed.
And for higher-risk audit engagements, an engagement quality reviewer also performs an independent, objective review. This person is a partner or managing director who was not part of the engagement team and who is specifically responsible for evaluating the significant judgments and conclusions reached by the team, including whether firm policy on AI use was followed.
This standard, he said, is closer to back-tracing the information than a casual read-through with skepticism baked into every step of the process.
"Professional judgment exercised by an experienced professional is necessary to determine the extent to which such review is required," said Ard. "The professional is expected to be able to defend the content to a regulator or in court exactly as if they had drafted it themselves without AI assistance."
Accountability and quality control
While the controls are meticulous, no system is perfect, so accountability is also a key piece of the puzzle. Overall, though, it works largely the same as it would for anything else. There was unanimous agreement that if AI generates false information that reaches an external audience, the AI is not to blame, no more than a hammer is to blame for a poor construction job.
"At the end of the day, and this is not rocket science, we are responsible and accountable to any work product we create, and so I almost don't differentiate between an outcome that is AI-created or human-created," said Armanino. "In the same way that we own our deliverables that humans create, we own our deliverables that are AI assisted. For us, it would never be acceptable if we release any type of work product that has errors and we catch it and we say, 'Hey, what happened here?' and the answer is, 'Oh, sorry, the AI said that.' For us, that's unacceptable."
Desai raised a similar point, saying that someone from Schellman has to approve the output and sign off on the work, representing that the information was accurate. While AI changes how the work is done, it does not change this fundamental accountability structure. She pointed out that it's not just the individual but the organization that is accountable for the work it produces, and accountability can involve not just the person who made the mistake but the firm as a whole.
"From an engineering perspective, you may improve prompts, retrieval techniques, guardrails or even change models. From an organizational perspective, you improve reviews, training and governance," she said.
But even at this scale, according to Armanino, the organization would respond much the same way as it would if a human had made a mistake with a client deliverable.
"I think we would evaluate any of these use cases in line with our core values. At the end of the day, if you make an error, whether it's AI-driven or not, we behave the same way. And that means being ultra-transparent with whoever was on the end of our work product, ultra-transparent with a client, and we'd make it right," he said.
Ard made a similar point, saying Forvis Mazars' guidance states directly that accountability remains with the professional, not the tool, and that using AI does not transfer accountability or professional responsibility. But he noted, like others, that it is unrealistic to expect a perfect result every time, so instead they focus on doing every single thing they can to prevent a mistake while acknowledging that something might still slip through.
"The distinction I would draw is: Our tolerance for skipping the review step is zero; our tolerance for occasional human review imperfection, like any quality system, is managed through training, supervision, engagement quality review and continuous improvement," he said.
Similarly, quality control procedures are generally the same regardless of whether or not AI was involved. Quality is quality and an error is an error, and a quality review program is meant to minimize those errors no matter where they came from. Hallucinations are just one more example of an error.
"A hallucination is a defect, and the purpose of quality management is to reduce defects to an acceptable level," said Shimamoto. "This should be addressed from a quality management lens. The 'quality triangle' expresses the key levers as scope, time and cost. Additionally, firms must consider the broader risks of the engagement. For example, reports are for public consumption (similar to public company financial statements), so the firm should consider potential reputational risk if there was a material defect in their report. This risk would be less prevalent if it was a report intended for internal management use."
Ard said the foundations of quality control do not change because of AI: His firm is explicit that the use of AI does not modify, override or substitute for any firm quality management control or professional standard, and AI-assisted content is subject to exactly the same supervision and engagement quality review requirements as manually prepared work. However, there are some slight differences when AI is involved.
"What is different is that we've layered AI-specific requirements on top of that existing system: a consideration of planned AI use and skepticism at the engagement planning stage, explicit consideration that AI outputs were reviewed and validated, and detailed guidance describing the specific ways AI-related risk shows up — hallucination patterns, overconfidence, automation bias, confirmation bias and anchoring bias — that our people are trained to watch for," he said. "Those are risks that don't have a clean human-error analogue, so we built new guidance to name them specifically, layered on top of, not replacing, our existing controls."
Reyes raised a similar point in that the black box nature of AI means firms can't afford to completely ignore the differences even if the core of quality management remains the same. He noted that if a person makes a mistake, someone can talk to them about their reasoning and trace how the mistake was made.
"If I was a partner in the firm, if I challenged the junior, 'Why did you say this in your working paper? Show me the trail of documents you gathered from someone to make you conclude that this is the issue for this client.' I'm going to make him trace his thought pattern to get that conclusion, something I cannot do with an LLM because an LLM is just so dynamically complex in how it gets to its conclusion. I need a mechanism where I can be comfortable with how it got to that point where I can exercise my professional judgment," he said.
Real talk on fake data
While hallucination risk will likely remain a constant so long as AI remains in existence, and while firms cannot say 100% they will never let hallucinated content get past their checks, they are far from helpless. Desai said that one of the biggest mistakes a firm can make in this regard, though, is in treating this as an IT project versus a governance and organizational one. Minimizing hallucination risk is less about the technology and more about the processes.
However, for these processes to work as intended, firms need to take a methodical, deliberate approach that emphasizes real human review versus just scanning a document and calling it a day, and having these processes baked into the heart of the workflow.
Training was also mentioned repeatedly as a vital measure in hallucination control, as people need to be educated as to what AI models do and why it is not the best idea to rely on them too much, perhaps by identifying the specific cognitive failure modes — described by Ard as overconfidence, automation bias, confirmation bias and anchoring — versus a generic "be careful with AI" message. In this regard, he said, firms also need to be specific about what they mean by "review" and ensure this definition is enforced with real measures that people will follow. And firms, he added, should also map their controls across every service line and practice area, rather than assuming a policy written for one part of the business protects all of it, and then close any gaps deliberately. This is likely sound advice for governance in general, including for AI.
While some may grumble at such stringent measures, Desai said they can actually be a competitive advantage. Done well, good governance won't stand in the way of a firm's AI ambitions, it will accelerate them.
"I think the biggest misconception is that AI governance is about how it is going to slow down AI and stifle innovation," she said. "To me, it's actually the opposite. Good governance lets you adopt AI faster because you understand the risks and have confidence in how it's being used. Also, it will be cost beneficial if you have to redo anything!"
Besides, she added, you don't want to see the alternatives.
"It can be much bigger than embarrassment," she said. "Obviously you have reputational damage and client trust. But depending on the engagement, you could also have regulatory scrutiny, litigation, contractual liability, professional discipline, insurance implications, and in some industries significant financial penalties. For professional services firms, trust is literally the product you're selling. If clients stop believing your work is reliable, that's much harder to recover from than fixing a technical issue."






