5 min read

GLM-5.2 nears the frontier, but safety lags

Z.ai’s GLM-5.2 nears frontier models on cyber and bio capabilities, but SaferAI found it refused none of its dangerous tasks.

Image: TechCrunch

Z.ai’s GLM-5.2 is only months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 on cyber and biological capabilities, according to a new evaluation from AI safety nonprofit SaferAI. But the same report highlights a sharper divide in safety behavior: GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it received through Z.ai’s public API.

The findings arrive as policymakers debate how to govern increasingly capable systems, including OpenAI’s GPT-5.6 Sol and Anthropic’s Mythos. They also intensify an argument already surrounding open-weight models: capability may be spreading faster than enforceable safeguards.

GLM-5.2's capability and safety gap

SaferAI found that Claude Opus 4.7 refused tasks so consistently that it could not complete CyberGym, a cybersecurity benchmark. GLM-5.2, by contrast, completed the offensive cyber and dual-use biology requests included in the evaluation without refusing any of them.

Recommended reading

A $10 ESP32 can run a tiny language model

“The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly,”

Henry Papadatos, executive director of SaferAI

The evaluation was conducted through Z.ai’s public API, where the company could apply controls. Those protections would not carry over when users download and run GLM-5.2's weights on their own hardware. A user could remove or alter safeguards, fine-tune the model, or change its system prompts.

That distinction is central to the open-weight debate. Closed-model providers such as OpenAI and Anthropic can use classifiers, refusal training, API controls, selective capability restrictions, pre-deployment evaluations, and risk assessments. None of those controls can be enforced in the same way once model weights are freely available.

The industry’s existing defenses are not perfect, either. Far.ai found hundreds of universal jailbreaks—reusable attack sequences that work against most harmful requests—in models including xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. The attacks combine techniques such as roleplaying, impersonating authority, fabricating conversation history, and repeated follow-up prompts.

Filtering training data is not a complete solution

Papadatos pointed to pre-training data filtering as one possible mitigation. The technique removes offensive cybersecurity material from the data used to train a model. Research cited by SaferAI suggests similar filtering can reduce hazardous biological knowledge without substantially damaging general performance.

Cybersecurity is harder. A model trained to be strong at coding will likely retain capabilities that can be used for hacking, making it difficult to separate commercially valuable programming ability from offensive cyber assistance. Because coding has become AI’s biggest source of revenue, developers have strong incentives to keep improving it.

Frontier labs have therefore explored narrower restrictions. Anthropic’s Opus 5, for example, can search for vulnerabilities in uncompiled source code but not compiled software, according to its system card. The stated logic is that limiting access to compiled targets makes offensive use more difficult.

Other options include withholding model weights when a system is judged too dangerous, publishing risk assessments, and requiring rigorous safety testing before release. SaferAI said Z.ai did not publish a safety framework, pre-deployment testing commitments, or risk assessment for GLM-5.2. TechCrunch said Z.ai did not respond when asked whether it had conducted internal or third-party frontier safety evaluations.

China’s regulatory model leaves questions open

Chinese officials have increasingly acknowledged the risks of advanced AI. At last month’s World AI Conference, President Xi Jinping emphasized the importance of open-weight models while also saying AI must remain under strict human control.

Graham Webster, who studies Chinese AI policy at Stanford’s Cyber Policy Center, said China has extensive AI regulations, but they have historically focused on politically sensitive content, misinformation, and social stability rather than catastrophic risks such as offensive cyber operations or biological misuse.

“U.S. AI thinkers are, in general, more concerned with this existential catastrophic [idea] than the Chinese community.”

Graham Webster, Stanford Cyber Policy Center

Webster said many Chinese policy researchers believe American companies would encounter a genuinely novel frontier risk first. He also argued that China’s real-name internet system gives authorities confidence that companies and users can be held accountable for how these technologies are used.

That model could potentially be extended to technical misuse, Webster said: the mechanisms used to make models refuse politically sensitive requests might be adapted to block offensive cyberattacks or harmful biological engineering instructions. But Chinese companies often coordinate with regulators privately, making it difficult to determine what testing took place before release.

The argument has also gained urgency after Hugging Face used GLM-5.2 to defend against an OpenAI breach. Clem Delangue, Hugging Face’s CEO, said the same systems that help stop an AI-powered attack could defend against millions of attacks and uncover vulnerabilities before criminals exploit them.

“The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them.”

Clem Delangue, CEO of Hugging Face

Papadatos said that defensive value is frequently overstated and does not justify making dangerous capabilities universally accessible. The report’s conclusion is difficult to avoid: GLM-5.2's progress makes open-weight systems more competitive, but its lack of refusal behavior—and the absence of a published safety framework or risk assessment—shows why capability parity is not the same as responsible release.

The practical imbalance is stark: attackers can change ransomware methods in a week, while a hospital cannot.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via TechCrunch

/ Keep reading