To reduce mistakes made by artificial intelligence (AI), companies have two main options: they can use a more reliable AI model or invest in verifying the responses it provides. The choice between these options depends on how important the AI's role is in their operations. For example, a poorly translated internal meeting used in marketing may not be as serious as incorrect legal advice, which could lead to significant consequences.
In practice, any company that uses a conversational AI tool will eventually face the same issue: the tool gives a wrong answer but presents it confidently. This phenomenon is known as a "hallucination," where the AI creates information that seems plausible but is not accurate. The immediate reaction is often to switch to a newer, more expensive, and supposedly more reliable AI model. Alternatively, companies might keep the cheaper model and have a human review its responses. Both approaches aim to reduce the number of errors reaching customers, but both come with costs. A recent study explores how companies can weigh these different costs.
This trade-off is already a real challenge. The Bank of France notes that many French companies are using these AI tools, even though the productivity gains are still limited. It is important to note that even a more reliable AI model does not eliminate the need for human verification. A 2025 study tested professional legal AI tools that claimed to avoid hallucinations, but between 17% and 33% of their responses still contained false information or incorrectly attributed statements to sources that did not support them.
Verification can take various forms, such as connecting the AI to a database of verified sources, generating multiple responses and selecting only the consistent ones, or having a staff member review the output. None of these methods is free, as they all require computing power or human labor. Our model compares the cost of using a paid, reliable AI model with the cost of verifying responses internally. In both cases, the goal is the same: ensuring reliability.
The decision of which option is cheaper depends on the nature of the AI's use. If a company frequently uses the AI for high-stakes tasks, such as legal advice or health-related decisions, paying more for a more reliable model may be more cost-effective than relying on thorough internal verification. The key factor is the risk of error and how much it could cost the company. Two companies using the same AI tool might choose different approaches based on the specific risks they face. For instance, using the AI for customer service carries less risk than using it for compliance or legal matters.
It is also important to clarify what constitutes a "high-stakes" use. The term covers two different factors: the benefit of a correct response and the cost of a wrong one. The cost of a wrong response, known as the "aversion to hallucinations," is not necessarily related to the benefit of a correct one. For example, a poorly worded marketing message can be corrected, but a wrong legal response could be expensive. High-stakes uses include legal advice, health-related decisions, and credit assessments, while lower-stakes uses might include internal translations or meeting summaries.
Finally, the question of who performs the verification remains. Some companies buy ready-made verified AI services, while others develop their own tools internally. A study on 3,143 companies found that many European companies use both approaches. When verification is outsourced, the cost includes not only the AI access and the verification work but also the vendor's profit margin. This margin is not just an extra cost; reducing it can be seen as buying a cheaper service with less thorough verification.
If a company chooses to handle verification internally, it still needs to manage and monitor the process. Simply moving the verification work inside the company does not guarantee its quality. The verification effort still exists, just in a different form. If the company detects an error, it can penalize its vendor. However, if it fails to notice an error, the vendor might reduce its efforts, knowing that the client might not detect the mistake. Our model focuses only on the errors that cost the company directly, but a wrong response can also harm others, such as customers or job candidates. It remains to be seen who will bear these additional costs.
Companies Weigh Costs of AI Reliability and Verification
AI-rewritten from original reportingHow it works
aiverificationhallucinationcost-benefitreliabilityrisk-management



