Back to blog

Open-weight AI models catch up to the frontier. The safety gap remains.

Open-weight AI models catch up to the frontier. The safety gap remains.

Open-weight AI models catch up to the frontier. The safety gap remains.

Open-weight AI models are progressing at a dizzying pace. The latest report from SaferAI, a nonprofit focused on AI safety, shows that GLM-5.2, the Chinese open-weight model from Z.ai, is only a few months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 — on cyber and bio capabilities, the two most sensitive domains.

But what worries experts is not so much the race for capabilities: it’s the growing divide between capabilities and safety practices.

What SaferAI’s evaluation reveals

SaferAI tested GLM-5.2 through Z.ai’s public API, on benchmarks designed to measure offensive risks:

  • Refusal of offensive cyber tasks: none. GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given.
  • Striking comparison: Claude Opus 4.7 refused so consistently that SaferAI could not complete CyberGym (a cyber-risk benchmark) on the model at all.

“The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly.” — Henry Papadatos, executive director of SaferAI

The fundamental problem with open weights

Z.ai could apply safety measures to its hosted API. But those protections become unenforceable once someone downloads the weights and runs them on their own hardware, where safeguards can be removed or modified, models fine-tuned, or system prompts changed.

Frontier developers like OpenAI and Anthropic rely on classifiers, refusal training, and API-level controls. These measures are far from foolproof — universal jailbreaks routinely bypass protections on deployed models, as Far.ai demonstrated on Grok 4.5 and Google DeepMind’s models.

The crucial point is simple: the safeguards built for closed models don’t work at all on open-weight models, which are designed to run on any infrastructure with any set of safeguards — or none.

Mitigation paths

Papadatos identifies several approaches:

  1. Pre-training data filtering: remove offensive cybersecurity information from training data and train on the curated dataset. Research suggests this can reduce hazardous biological knowledge without harming overall performance. For cybersecurity, however, filtering is much less practical: it’s difficult to train a model that excels at coding without it also being a good hacker.

  2. Selective restriction: limit the kinds of cybersecurity assistance models provide. Claude Opus 5, for example, can search for vulnerabilities in uncompiled source code, but not in compiled code.

  3. Pre-deployment evaluations: rigorous safety evaluations, published risk assessments, and withholding model weights if a system is perceived as too dangerous.

In GLM-5.2’s case, SaferAI notes that Z.ai published no safety framework, no pre-deployment testing commitments, and no risk assessment for the model.

The Chinese position

Chinese leaders have increasingly acknowledged the risks of advanced AI. At the World AI Conference last month, President Xi Jinping emphasized the importance of open-weight models while stressing the necessity of keeping AI under strict human control.

Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, adds nuance: China has robust AI regulations, but historically focused on politically sensitive content, misinformation, and social stability — not catastrophic AI risks. Chinese policy thinkers, he notes, are generally less concerned with existential risk than their American counterparts.

The open-weight defense argument

Open-weight advocates defend releasing the weights for one crucial reason: defense. Hugging Face relied on GLM-5.2 to defend itself against the OpenAI breach.

“The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them.” — Clem Delangue, CEO of Hugging Face

Papadatos says that benefit is often overstated — and doesn’t mean we should open-source dangerous capabilities:

“The main point in my mind is that we shouldn’t just accept that dangerous capabilities are easily accessible by anyone anywhere.”

Key takeaways

The open-weight vs safety debate is not binary. It’s about balancing the spread of useful capabilities against the prevention of dangerous uses — with one clear finding: by default, attackers adopt new tools faster than defenders publish their safeguards.

For businesses and developers, this means that integrating open-weight models into production requires a case-specific safety assessment: model provenance, mitigation measures, exposure surfaces, and the ability to monitor usage. Technology is not neutral — and its safety is not declared, it’s evaluated.

Already using AI in your business, or planning to? Our team helps companies adopt AI in a secure and strategic way — not just a performant one. Get a free AI maturity diagnostic and walk away with a concrete action plan tailored to your needs and risks.

Have a similar project?

Get a free diagnostic of your online presence and personalized recommendations.

Free Diagnostic

Don't leave without your gift!

Download our free "Digital Diagnostic" guide to discover how to improve your online presence.