Skip to main content
A person's hand interacting with a digital interface, illustrating the open-weight AI safety gap and removable protections.

Editorial illustration for Open-Weight AI Safety Gap Persists as Users Can Remove Hosted Protections

Open-Weight AI Models Lack Safety Protections

Open-Weight AI Safety Gap Persists as Users Can Remove Hosted Protections

4 min read

GLM-5.2, the open-weight model released by China's Z.ai, refused zero offensive cyber tasks and zero dual-use biology tasks in a new evaluation from AI safety nonprofit SaferAI. Claude Opus 4.7, tested for comparison, refused so often that SaferAI couldn't even finish running CyberGym, the cybersecurity benchmark OpenAI used before last month's Hugging Face breach, on it. The capability gap between GLM-5.2 and models like GPT-5.5 and Claude Opus 4.7 is now measured in months, not years, according to SaferAI's findings. The safety gap, by contrast, is widening.

That split matters because open-weight models can't be recalled once released. Anyone can download the weights, strip out guardrails, and run the model on their own hardware, out of reach of the usage policies and monitoring that companies like OpenAI and Anthropic build around their hosted systems. As GPT-5.6 Sol, Anthropic's Mythos, and Chinese competitors like GLM-5.2 push the frontier of what AI systems can do in cybersecurity and biology, policymakers are confronting a harder question than raw capability: what happens when that capability ships with no way to police how it's used.

According to SaferAI’s evaluation, which the nonprofit ran via Z.ai’s public API, GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given.

Why this matters

For anyone building on open-weight models, the SaferAI report is a reminder that capability parity doesn't mean safety parity. GLM-5.2 sitting just months behind GPT-5.5 and Claude Opus 4.7 on cyber and bio benchmarks sounds like a win for open ecosystems, and in some ways it is. But once weights are downloadable, every safeguard Z.ai builds into its hosted API is a suggestion, not a control.

Fine-tuning, stripped system prompts, local deployment: all of it sits outside the developer's reach the moment someone else has the file. That's a different risk profile than what OpenAI and Anthropic manage through API-only access, and policymakers debating governance frameworks right now are largely writing rules for the latter. If you're evaluating GLM-5.2 or any open-weight model for production use, the benchmark scores are only half the story.

The real question is who's running the weights downstream, what they've changed, and whether you'd even know. Capability convergence without safety convergence isn't a footnote here. It's the whole gap the industry still hasn't closed.

Common Questions Answered

Why did GLM-5.2 perform worse than Claude Opus 4.7 on SaferAI's safety evaluation?

According to SaferAI's evaluation, GLM-5.2 refused zero offensive cyber tasks and zero dual-use biology tasks, meaning it failed to decline harmful requests. In contrast, Claude Opus 4.7 refused requests so frequently that SaferAI couldn't even complete running the CyberGym cybersecurity benchmark on it, demonstrating superior safety guardrails.

What is the capability gap between open-weight models like GLM-5.2 and frontier models?

The capability gap between GLM-5.2 and models like GPT-5.5 and Claude Opus 4.7 is now measured in months rather than years, indicating that open-weight models are rapidly catching up to frontier closed-source models. This narrowing gap represents significant progress in the open-source AI ecosystem.

How can users bypass the safety protections built into GLM-5.2's hosted API?

Once the model weights are downloadable, users can circumvent Z.ai's hosted safeguards through fine-tuning, stripping system prompts, or deploying the model locally. These methods allow developers to modify the model outside of the developer's control, making any safety measures built into the hosted API merely suggestions rather than enforced controls.

What does SaferAI's report reveal about the relationship between capability parity and safety parity in open-weight models?

SaferAI's evaluation demonstrates that capability parity does not automatically mean safety parity in open-weight models. While GLM-5.2 achieving near-parity with frontier models on cyber and bio benchmarks represents progress for open ecosystems, the downloadable weights create a fundamental safety vulnerability that cannot be enforced once the model is deployed locally.

LIVE22:14Anthropic Secures USD 10 Billion AI Cloud Deal With Nvidia Partner Volta