Gabriela Moreira, CEO of Quint, highlights the growing risks of using AI tools to write code and emphasizes the need for early checks and processes to prevent hard-to-maintain code and critical bugs. She explains that AI, while powerful, should be a tool that humans control and guide, rather than a replacement for human judgment. One major concern is that AI can generate a large amount of code that may contain errors, and the complexity of modern systems can make it hard for humans to fully understand or track all the elements of what the AI produces. As a result, traditional practices like code review and monitoring become more challenging at scale. Systems that are especially difficult to test—such as those that run simultaneously across multiple processes or handle unpredictable data—pose the greatest risk when using AI-generated code. These systems have many possible behaviors, making it harder to detect flaws before they cause problems. AI is a generation tool, and its effectiveness depends on humans' ability to judge whether the output is correct. If a system is hard to test or verify, it can create many "blind spots" where errors go unnoticed. Deploying AI-generated code without proper verification can introduce hidden risks, especially in terms of security and operations. In this scenario, defenders must secure all parts of their code, while attackers only need to find a single vulnerability to cause significant damage. Formal verification, a rigorous method for proving code correctness, is one of the strongest ways to address this imbalance, but it's not yet practical for verifying all code. As a result, vulnerabilities can still exist in unverified components. Moreira advises against relying on a single solution and instead recommends combining multiple strategies to manage these risks effectively. Identifying design-level bugs early in the development process is crucial because it’s much easier to detect and fix them when the system is still in a manageable state. Moreira is particularly concerned about the current trend of relying on markdown files for design, which can't be compiled or executed and may not reflect real-world behavior. To ensure AI-generated code is reliable before it goes into production, development teams should use tests that provide real assurance, not just automated tests that make arbitrary checks and require constant updates. Tests that focus on specific behaviors and prevent regression are more effective. Mutation testing, which evaluates how well tests catch errors, and formal models, which can either generate test scenarios or validate existing tests, are also valuable tools. The amount of human oversight needed to balance speed and reliability depends on the specific context and risk level. Moreira suggests that the best approach is to implement processes that allow teams to confidently deploy new code without fear of breaking existing functionality. The worst-case scenario is a security breach or data leak, while the second worst is when a system becomes so unstable that any attempt to fix it makes things worse. Moreira believes AI-generated code is especially prone to this kind of instability. However, she notes that the way out of this situation is similar to how developers historically handled legacy code—by first adding enough tests to ensure confidence in the system, and then making changes carefully and incrementally.