Open-weight AI models are advancing at a dizzying pace. The latest report from SaferAI, a nonprofit organization specializing in AI safety, shows that GLM-5.2, the open-weight Chinese model from Z.ai, is now only months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 — in cyber and bio capabilities, the two most sensitive domains.
But what worries experts isn’t so much the capabilities race: it’s the growing gap between capabilities and security practices.
What SaferAI’s evaluation reveals
SaferAI tested GLM-5.2 via Z.ai’s public API, on benchmarks designed to measure offensive risks:
- No refusal of offensive cyber tasks whatsoever. GLM-5.2 refused none of the offensive cyber or dual-use biology tasks submitted to it.
- Striking comparison: Claude Opus 4.7 refused so systematically that SaferAI couldn’t even complete CyberGym (a cyber risk benchmark) on the model.
“The capability frontier is not the risk frontier. Mitigation measures must also be taken into account to properly assess risk.” — Henry Papadatos, Executive Director of SaferAI
The fundamental problem with open-weight models
Z.ai could apply safety measures on its hosted API. But these protections become inapplicable as soon as someone downloads the weights and runs them on their own hardware: safeguards can then be removed, models modified through fine-tuning, or system prompts changed.
Closed model developers (OpenAI, Anthropic) rely on classifiers, refusal training, and API-level controls. These measures are far from foolproof — universal jailbreaks regularly succeed in bypassing protections of deployed models, as shown by Far.ai on Grok 4.5 or Google DeepMind models.
But the crucial point is simple: safeguards designed for closed models don’t work at all on open-weight models, which are designed to run on any infrastructure, with any set of protections — or none.
Mitigation pathways
Papadatos identifies several possible approaches:
-
Training data filtering: remove offensive cybersecurity information from training data, then train the model on the cleaned corpus. Research suggests this can reduce dangerous biological knowledge without degrading general performance. However, for cybersecurity, filtering is far less practical: it’s difficult to train a model that excels in code without it also being a good hacker.
-
Selective restriction: limit the type of cyber assistance models provide. Claude Opus 5, for example, can search for vulnerabilities in uncompiled source code but not in compiled code.
-
Pre-deployment evaluations: rigorous safety evaluations, risk analysis publication, and weight retention if a system is deemed too dangerous.
In the case of GLM-5.2, SaferAI highlights that Z.ai published no safety framework, no pre-deployment testing commitments, and no risk assessment for the model.
The Chinese position
Chinese leaders are increasingly acknowledging the risks of advanced AI. At the World AI Conference last month, Xi Jinping emphasized the importance of open-weight models while stressing the need to ensure AI remains a tool under strict human control.
Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, offers nuance: China has robust regulations, but historically focused on politically sensitive content, disinformation, and social stability — not catastrophic AI risks. Chinese thinkers, he adds, are generally less concerned about existential risk than their American counterparts.
The counterargument from open-weight advocates
Supporters of open-weight defend releasing weights for a crucial reason: defense. Hugging Face relied on GLM-5.2 to defend against the OpenAI breach.
“The same systems that helped stop an AI-driven cyberattack can now help defend against millions of cyberattacks every day, and identify vulnerabilities before attackers exploit them.” — Clem Delangue, CEO of Hugging Face
Papadatos believes this advantage is often overstated — and doesn’t justify open-sourcing dangerous capabilities:
“The main point, in my view, is that we shouldn’t simply accept that dangerous capabilities are easily accessible to anyone, anywhere.”
Key takeaways
The open-weight vs. security debate is not binary. It’s about balancing the diffusion of useful capabilities with the prevention of dangerous uses — with a clear observation: by default, attackers adopt new tools faster than defenders publish their safeguards.
For businesses and developers, this means integrating open-weight models into production requires a security evaluation specific to each use case: model provenance, mitigation measures, exposure surfaces, and the ability to monitor usage. Technology is not neutral — and its security isn’t decreed, it’s assessed.
Already using AI in your business, or planning to? Our team helps organizations integrate AI in a secure and strategic way — not just a performant one. Get a free audit of your AI maturity and leave with a concrete action plan, tailored to your needs and your risks.
