In early 2025, the Federal Trade Commission (FTC) finalized a $193,000 agreement with DoNotPay, a startup that had marketed itself as the world’s first robot lawyer. According to the FTC, DoNotPay had not demonstrated that its results were comparable to those of a licensed attorney. The company had also faced a class action lawsuit shortly after its launch for offering legal services in California without proper licensing. This case highlights a growing concern about the gap between the bold claims made by large AI companies and the reality of their performance. Legal experts are now grappling with a lack of precedent for handling the complexities of AI-related failures, creating uncertainty in areas like corporate responsibility and regulatory compliance. As the debate continues, a new category of AI software is emerging, designed with a clear chain of responsibility in mind. These systems aim to address the legal and ethical concerns that have arisen from the rapid development of AI technologies. ZDNET spoke with four experts in the field to explore the underlying technology and assess whether the security promises of these systems hold up in practice. One such design model gaining traction is Human-in-the-loop (HITL), in which AI systems submit complex decisions for human review before executing them. This approach is intended to ensure rigorous decision-making, especially in sensitive areas like health operations, financial transactions, and legal decisions. However, the term "HITL" is used flexibly across different systems, and not all implementations involve real-time human oversight during critical tasks. Experts warn that the effectiveness of HITL systems depends heavily on how they handle uncertainty or errors in AI models. Akash Thakur, an AI reliability engineer, explains that the main issue is not the model itself but how systems respond to potential failures. Rather than treating AI errors as worst-case scenarios, HITL systems are designed to act as real-time audit tools, intercepting decisions that are likely to be incorrect. However, many systems rely too heavily on the AI’s confidence score to trigger a human review, which may not accurately reflect the actual accuracy of the result. Daniel Gamber, CEO of Cambrion, points out that confidence scores only indicate the model’s ability to process information, not the correctness of the outcome. Beyond the technical aspects, the implementation of HITL systems also involves addressing biases and ensuring audit trails. Asim Husain, co-founder of Alterion, emphasizes that the response to AI errors must be tailored to the organization’s risk tolerance and that the system must know to whom to escalate issues. Additionally, the European Union’s AI Act mandates verifiable supervision as a key compliance requirement. Training human reviewers to critically assess AI outputs is crucial, but biases in the AI model itself can undermine even the best review protocols. These challenges underscore the need for robust, adaptable systems that can withstand scrutiny and ensure accountability in the evolving AI landscape.