
An open-weight AI model from Chinese developer Z.ai has closed the performance gap with top proprietary systems from OpenAI and Anthropic on offensive cyber and biological benchmarks, while diverging sharply on safety, according to a new report from nonprofit SaferAI.
The report assessed GLM-5.2, accessed via Z.ai’s public API, and found it refused none of the offensive cyber or biology tasks in the evaluation. In contrast, Anthropic’s Claude Opus 4.7 “refused so consistently that SaferAI could not complete CyberGym on it at all,” the report states. CyberGym is a cybersecurity capability benchmark that OpenAI used before last month’s Hugging Face breach.
SaferAI’s findings reanimate longstanding warnings that open-weight releases grant bad actors unrestricted access to advanced AI capabilities. Once model weights are downloaded, users can strip safety filters, fine-tune the model on harmful data, or alter system prompts, rendering any upstream safeguards unenforceable.
“The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly,” Henry Papadatos, executive director of SaferAI, told TechCrunch.
Frontier developers like OpenAI and Anthropic embed safeguards such as classifiers, refusal training, and API-level controls. Yet even these are routinely bypassed by jailbreaks. The nonprofit Far.ai recently documented hundreds of universal jailbreaks—reusable attack sequences—against models including xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. Attackers combine roleplaying, authority impersonation, fake conversation history, and follow-up prompts to amplify vulnerabilities.
For open-weight models, even those imperfect protections evaporate when the model runs on user-controlled hardware. Papadatos advocates designing models so that “the good capabilities — the safe ones — are accessible to anyone, and then we try to remove the bad ones, even in an open source fashion.”
One proposed technique, pre-training data filtering, removes offensive cybersecurity knowledge from training sets before model training. Research suggests this can reduce hazardous biological knowledge without degrading overall performance, but for cybersecurity the approach is less practical. Coding is AI’s largest commercial application, and it is hard to train a model that excels at legitimate programming but cannot assist with hacking. Frontier labs have instead relied on releasing models that selectively restrict assistance—for example, Anthropic’s Opus 5 can search for vulnerabilities in uncompiled source code but not compiled software, the company’s system card states. Other mitigations include rigorous pre-release safety evaluations, public risk assessments, and withholding weights entirely for models deemed too dangerous.
Z.ai, however, did not publish a safety framework, pre-deployment testing commitments, or a risk assessment for GLM-5.2, according to SaferAI. TechCrunch requested comment on whether Z.ai conducted internal or third-party frontier safety evaluations before release but received no response.
The release comes as Chinese policymakers increasingly discuss AI risks. At last month’s World AI Conference, President Xi Jinping stressed the importance of open-weight models while emphasizing the need to keep AI under strict human control. Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, told TechCrunch that China’s AI regulations have historically targeted politically sensitive content, misinformation, and social stability, not catastrophic risks such as offensive cyber capabilities or biological misuse.
“U.S. AI thinkers are, in general, more concerned with this existential catastrophic [idea] than the Chinese community,” Webster said, adding that many Chinese policy researchers believe American companies will encounter any novel frontier risks first. Webster noted that China’s real-name internet system and the government’s ability to hold companies and users accountable provide a mechanism for control. Companies could potentially adapt content-refusal filters designed for political topics to block offensive cyber or biological assistance. The opaque coordination between Chinese firms and regulators makes it difficult to know what pre-release testing occurs.
Advocates of open-weight models argue that weight release strengthens collective cybersecurity. Hugging Face CEO Clem Delangue, posting on social media this week, said open models “helped stop an AI-powered cyberattack” and now “help defend against millions of cyberattacks every day.” SaferAI’s Papadatos considers that benefit overstated and cautioned against releasing dangerous capabilities. “We shouldn’t just accept that dangerous capabilities are easily accessible by anyone anywhere,” he said. He emphasized the asymmetric adoption speed: “a ransomware group can change its methods in a week. A hospital cannot.”
See an error? Read our corrections policy or email [email protected].
TECHNOMALIST

