Three cybersecurity researchers in India used a competing AI model from Anthropic to gain unauthorized access to OpenAI, highlighting the potential risks that advanced artificial intelligence poses to companies and their users. Mohan Pedhapati, one of the researchers from a company called Hacktron, explained that more advanced AI models make hacking easier for experienced hackers, increasing the likelihood that criminals could do the same. He described each new AI model as a "force multiplier" that significantly enhances hacking capabilities.
In a report released last week, Hacktron detailed how they used the Claude AI model in late July to infiltrate the OpenAI community forum. This forum is a space where users of ChatGPT and Codex, OpenAI's coding assistant, can ask questions about the products. Users log in with their OpenAI accounts. The researchers used a version of the Claude model called Opus 4.8 to find a flaw in the code of a third-party service called Discourse, which is used to run the forum. However, they were unable to exploit the flaw with that model. The next day, a more advanced version, Claude Opus 5, was released. Using this updated model, the Hacktron team was able to access the ChatGPT and Codex accounts of some OpenAI users who were on the forum.
Once they had access to these accounts, the researchers said they could have also accessed other apps connected to the accounts, such as email and Slack. They could also view the conversations users had on ChatGPT. Some of these users were OpenAI employees, whose accounts were linked to internal email systems. The Wall Street Journal first reported on this incident.
Hacktron's actions were part of a common practice in the tech industry known as "bug bounties," where companies pay outside researchers for discovering and reporting security vulnerabilities. Hacktron was attempting to help OpenAI by identifying weaknesses in their system and then gaining access to a more powerful AI model with fewer restrictions. Pedhapati said that without the help of Claude, the hack would have taken him two to three months. But with the AI model, the entire operation was completed in less than three days. He estimates that with newer models, such a hack could now take less than a day.
Pedhapati and his team have previously hacked into companies like Apple, Google, Facebook, Discord, and Microsoft Teams. He expressed concern that companies like OpenAI are not sufficiently securing their systems given the power of their AI models. He warned that companies must assume they could be hacked and build their security accordingly. He believes that many AI labs, including OpenAI and Anthropic, are moving too quickly without adequate safeguards.
OpenAI responded to the breach by thanking the researchers and stating that they had restricted access to affected accounts. In the wake of the Hugging Face hack and growing concerns about AI risks, companies and experts are discussing the broader societal implications of powerful AI systems. One of the main concerns is the potential for these models to be stolen, particularly the data that makes them so effective. Experts like Nicholas Leiserson from the Institute for Security and Technology argue that AI development has national security implications and that the pace of innovation needs to be carefully considered. Despite some calls for slowing down, leading AI labs have prioritized rapid development for years, leaving the determination of appropriate risk levels to the labs themselves.
Cybersecurity Researchers Exploit AI Models to Access OpenAI User Accounts
AI-rewritten from original reportingHow it works
ai-hackingcybersecurityopenaianthropicbug-bountyai-risks
Original sources:
- 🇺🇸CBS News



