Open-Weight AI Models Are Catching Up to the Frontier — But Not to Its Guardrails

Read in : AR EN FR

Open-weight AI models are advancing at a dizzying pace. The latest report from SaferAI, a nonprofit organization specializing in AI safety, shows that GLM-5.2, the open-weight Chinese model from Z.ai, is now only months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 — in cyber and bio capabilities, the two most sensitive domains.

But what worries experts isn’t so much the capabilities race: it’s the growing gap between capabilities and security practices.

What SaferAI’s evaluation reveals

SaferAI tested GLM-5.2 via Z.ai’s public API, on benchmarks designed to measure offensive risks:

“The capability frontier is not the risk frontier. Mitigation measures must also be taken into account to properly assess risk.” — Henry Papadatos, Executive Director of SaferAI

The fundamental problem with open-weight models

Z.ai could apply safety measures on its hosted API. But these protections become inapplicable as soon as someone downloads the weights and runs them on their own hardware: safeguards can then be removed, models modified through fine-tuning, or system prompts changed.

Closed model developers (OpenAI, Anthropic) rely on classifiers, refusal training, and API-level controls. These measures are far from foolproof — universal jailbreaks regularly succeed in bypassing protections of deployed models, as shown by Far.ai on Grok 4.5 or Google DeepMind models.

But the crucial point is simple: safeguards designed for closed models don’t work at all on open-weight models, which are designed to run on any infrastructure, with any set of protections — or none.

Mitigation pathways

Papadatos identifies several possible approaches:

  1. Training data filtering: remove offensive cybersecurity information from training data, then train the model on the cleaned corpus. Research suggests this can reduce dangerous biological knowledge without degrading general performance. However, for cybersecurity, filtering is far less practical: it’s difficult to train a model that excels in code without it also being a good hacker.

  2. Selective restriction: limit the type of cyber assistance models provide. Claude Opus 5, for example, can search for vulnerabilities in uncompiled source code but not in compiled code.

  3. Pre-deployment evaluations: rigorous safety evaluations, risk analysis publication, and weight retention if a system is deemed too dangerous.

In the case of GLM-5.2, SaferAI highlights that Z.ai published no safety framework, no pre-deployment testing commitments, and no risk assessment for the model.

The Chinese position

Chinese leaders are increasingly acknowledging the risks of advanced AI. At the World AI Conference last month, Xi Jinping emphasized the importance of open-weight models while stressing the need to ensure AI remains a tool under strict human control.

Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, offers nuance: China has robust regulations, but historically focused on politically sensitive content, disinformation, and social stability — not catastrophic AI risks. Chinese thinkers, he adds, are generally less concerned about existential risk than their American counterparts.

The counterargument from open-weight advocates

Supporters of open-weight defend releasing weights for a crucial reason: defense. Hugging Face relied on GLM-5.2 to defend against the OpenAI breach.

“The same systems that helped stop an AI-driven cyberattack can now help defend against millions of cyberattacks every day, and identify vulnerabilities before attackers exploit them.” — Clem Delangue, CEO of Hugging Face

Papadatos believes this advantage is often overstated — and doesn’t justify open-sourcing dangerous capabilities:

“The main point, in my view, is that we shouldn’t simply accept that dangerous capabilities are easily accessible to anyone, anywhere.”

Key takeaways

The open-weight vs. security debate is not binary. It’s about balancing the diffusion of useful capabilities with the prevention of dangerous uses — with a clear observation: by default, attackers adopt new tools faster than defenders publish their safeguards.

For businesses and developers, this means integrating open-weight models into production requires a security evaluation specific to each use case: model provenance, mitigation measures, exposure surfaces, and the ability to monitor usage. Technology is not neutral — and its security isn’t decreed, it’s assessed.

Already using AI in your business, or planning to? Our team helps organizations integrate AI in a secure and strategic way — not just a performant one. Get a free audit of your AI maturity and leave with a concrete action plan, tailored to your needs and your risks.

Izri.Online

Écrit par Izri.Online

Équipe Izri.Online — Agence digitale au Maroc. 2 Humains + 10 Agents IA.

Do you have a digital project?

Free and no-obligation diagnostic — we analyze your online presence and propose concrete solutions.

Request my diagnostic
0
0