Five months ago, the release of Claude Mythos Preview, an AI model capable of creating complex cyber exploits on its own, was announced. This development raised concerns that similar capabilities could soon be replicated by other AI models, making it easier for cybercriminals to launch powerful attacks. In response, Claude Mythos Preview was made available only through the Glasswing project, which allowed trusted cybersecurity experts to identify over 10,000 vulnerabilities in critical software before malicious actors could access similar AI tools. However, models with similar capabilities are now available to the public. This article examines GLM-5.3, a new AI model developed by Zhipu AI (known internationally as Z.ai), which, like Claude Mythos Preview, can autonomously create end-to-end computer exploits. Unlike Claude, however, GLM-5.3 lacks significant security measures to prevent its misuse. Tests showed that attackers could bypass its security in 64% to 100% of cases using simple techniques, while similar attacks failed against the Claude models, which have built-in protections.
On September 17, the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST) evaluated GLM-5.3 and found it to be the most advanced open-weight AI model in terms of cybersecurity capabilities published to date. However, it lags about four months behind the best American AI models in cybersecurity benchmarks. These findings align with other assessments, which note that American models are generally more secure, as they are only available to approved users, while GLM-5.3 is freely downloadable by anyone. This article explores how easily the security measures of GLM-5.3 can be bypassed or removed, highlighting the risks and opportunities this model presents.
To understand how GLM-5.3 might help malicious actors detect and exploit real software vulnerabilities, researchers tested it using automated benchmarks and workflows involving human input. The models were run in isolated environments, attacking only offline targets configured for the tests. The focus was on exploit development, where Claude Mythos Preview showed a notable improvement over earlier Claude models. On ExploitBench, a test measuring the ability of AI models to exploit known vulnerabilities in Google Chrome's V8 engine, GLM-5.3 succeeded in 50 out of 410 attempts, while Claude Mythos Preview achieved 56 out of 410. In another benchmark, Binary Exploitation, which tests the ability to detect and exploit vulnerabilities in popular open-source projects, GLM-5.3 achieved a complete control flow hijack in 4% of attempts, compared to 6% for Claude Mythos Preview. While GLM-5.3 performed slightly worse than Claude, it still represents a major leap forward, as earlier models like Claude Opus 4.6 and GLM-5.2 achieved zero success in these tests.
In another test, GLM-5.3 was used by human researchers to identify and exploit new vulnerabilities in software systems. In one experiment, a researcher used GLM-5.3 on a sandboxed machine with a popular web browser on Linux. Within a day, the model identified several previously unknown vulnerabilities in the browser's JavaScript engine and combined them into a functional exploit that could read arbitrary files from a user's computer. This exploit targets the Linux version of the browser, as that was the only environment available to the model, though similar vulnerabilities might affect other platforms. The researcher also used GLM-5.3 to find exploitable flaws in wireless and graphics drivers, as well as network-connected device software. These findings are currently being reviewed, and if necessary, they will be reported to the relevant maintenance teams. In a second session, a researcher used a lighter version of GLM-5.3, called GLM-5.3-Flash, to develop an exploit for a known vulnerability in Google Chrome. The model successfully transformed public patch details into an operational attack, bypassing a security feature called pointer authentication in about eight hours of processing time.
Chinese AI Model GLM-5.3 Demonstrates Advanced Cyber Exploit Capabilities with Weak Security Measures
AI-rewritten from original reportingHow it works
ai-cybersecurityglm-5-3open-weight-modelcyber-exploitssecurity-risks



