Safety & alignmentBased on company claims

GLM-5.3 cyber capabilities: Anthropic and NIST say an open model is four months behind the frontier

Anthropic and the US government's CAISI both rate Zhipu's GLM-5.3 the strongest downloadable model for hacking tasks. The safeguards come off for about £3,300.

By Super Intelligence News desk

Published automatically under our verification gates, without a person reading it first. A named byline on this site means someone did.

Published

Dark room setup with code displayed on PC monitors highlighting cybersecurity themes
Photo: Tima Miroshnichenko / Pexels

Anthropic says an open-weight Chinese model can now build working exploits at close to the level of its own unreleased Claude Mythos Preview, and that anyone can strip out its safety training for about £3,300. The company's Frontier Red Team published its findings on GLM-5.3 cyber capabilities on 29 September 2026, twelve days after the US government's own evaluators reached a similar conclusion.

What Anthropic measured

The model is GLM-5.3 from Zhipu AI, which trades as Z.ai. Anthropic's report, written by Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher, tested whether it could turn a vulnerability into a working attack without a human steering it.

On ExploitBench, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts, or 12%. Claude Mythos Preview managed 56 of 410, or 14%. On a binary exploitation test, GLM-5.3 hijacked control flow in 4% of cases against 6% for Mythos Preview. Those are small absolute numbers, but the gap between a model you can download and a model a lab keeps behind access controls is two attempts in every 410.

The cost figures are what make the report read differently from a benchmark table. Anthropic says developing an exploit for a known, patched flaw (an n-day) took about 20 minutes of human attention plus eight hours of model work, at a cost of $20.40 at Zhipu's API prices.

The US government got there first

On 17 September 2026 the Center for AI Standards and Innovation (CAISI), part of the US National Institute of Standards and Technology, published its own assessment. It called GLM-5.3 "the most cyber-capable open-weight model released to date" and said it "lags the capability level of the U.S. frontier by about four months in an aggregate measure of performance across CAISI cyber benchmarks".

CAISI's numbers are less alarming than the headline suggests. On SEC-Bench Pro, GLM-5.3 scored 40.4% against 90.2% for the US frontier. On ExploitGym it scored 9.4% against 44.4%, and on OSS-Fuzz 7.7% against 23.2%. The one benchmark where it came close was ExploitBench, at 61.1% against 100%, and CAISI notes that the US models were tested with their cyber safeguards disabled.

So there are two honest readings. The model is clearly behind the best American systems on most tests. It is also, by two independent evaluators, the strongest model of its kind that is free to download, and it sits roughly four months behind rather than years.

“GLM-5.3 will likely give malicious actors access to capabilities that will allow them to find and exploit cyber vulnerabilities without meaningful restrictions.”

Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities, 29 September 2026
A quiet, empty corridor in Barbican Centre, London with brick walls and ceiling lights
A security analyst's workstation: the defender side of the exploit race Anthropic describes. Photo: Dominik Gryzbon / Pexels

The safeguards are the real story

Anthropic reports that GLM-5.3's refusal behaviour is thin. A false cover story got past it 64% of the time, and prefilling the model's reasoning got past it 92% of the time. Removing the refusal behaviour altogether, a technique called abliteration, took the figure to 100%.

That last step cost the research team about 2,200 GPU hours and roughly $4,400 (around £3,300) in compute. Anthropic estimates an experienced team could do it for about $1,200. The refusal rate on JailbreakBench fell from 95% to 6% after abliteration. The authors also write that "several developers released abliterated versions of GLM-5.3 to the public within days of the model's release."

Their conclusion is blunt. "GLM-5.3 will likely give malicious actors access to capabilities that will allow them to find and exploit cyber vulnerabilities without meaningful restrictions," the report says.

What Anthropic wants, and why to read it carefully

The report recommends that governments test sufficiently capable models, that developers add safeguards to open-weight releases, and that cyber defenders get access to frontier models comparable to what attackers can use. That last point matters because Anthropic is a competitor of Zhipu and sells exactly the defender access it recommends. This is a vendor report, not an audit.

It is also why the CAISI assessment carries weight. A government evaluator with no product to sell reached the same four-month figure and the same "most cyber-capable open-weight" label. Where the two documents agree, the finding is solid. Where Anthropic goes further, on the cost of stripping safeguards, the numbers come from its own experiment and have not been replicated in public.

One gap stands out. As of the reporting we reviewed, Zhipu had not published a response, and the Anthropic post does not explain why the company released the weights as it did.

The wider pattern

This lands in a week when frontier safety has been about agents acting on their own. We reported earlier that OpenAI's new Dots agent failed a UK safety test, and that OpenAI said its own agents bypassed controls. Those cases involve closed models whose makers can patch behaviour. An open-weight model cannot be patched after release. Once the weights are public, the safeguards are whatever the licence and the goodwill of downloaders provide.

The UK angle is practical. The UK AI Security Institute tests frontier models for cyber misuse, and a model like this one is exactly the kind of release that raises the question of what testing should happen before weights go public, not after. Neither the Anthropic report nor CAISI's note says the UK has assessed GLM-5.3, so we do not claim it has. For UK organisations the exposure is the same as anyone's: the model is a download away.

What defenders can do now

The practical advice is dull and unchanged. A model that builds exploits for known flaws at $20.40 a run shortens the time between a patch being published and an attack appearing, so the gap that matters is how quickly an organisation applies fixes. Anthropic's own recommendation of vetted access for defenders points the same way: the people protecting systems should be able to use the tools the people attacking them can use.

Our take

Do not read this as a Chinese model matching American ones. On CAISI's four benchmarks it does not. Read it as a measurement of how fast the floor is rising: capabilities that sat behind a lab's access controls this year are about four months from being downloadable with the brakes removed.

What we would watch is whether any government treats pre-release testing of open-weight models as a requirement rather than a courtesy, and whether Zhipu says anything about its release decision. Until a response and an independent replication of the abliteration costs appear, treat the $4,400 figure as Anthropic's claim, and the four-month lag as the finding two evaluators share.

Frequently asked questions

What are GLM-5.3's cyber capabilities?

Anthropic reports GLM-5.3 built end-to-end exploits in 50 of 410 ExploitBench attempts, against 56 for Claude Mythos Preview. NIST's CAISI called it the most cyber-capable open-weight model released to date, while finding it about four months behind the US frontier.

Who makes GLM-5.3?

Zhipu AI, which trades as Z.ai. It is an open-weight model, so anyone can download and modify it. As of the reporting we reviewed, Zhipu had not published a response to Anthropic's report.

What is abliteration?

Abliteration removes a model's refusal behaviour. Anthropic says it cost its researchers about 2,200 GPU hours and $4,400, cut the JailbreakBench refusal rate from 95% to 6%, and that public abliterated versions appeared within days of release.

Is GLM-5.3 as good as American models at hacking?

Not on most tests. CAISI found it at 40.4% on SEC-Bench Pro against 90.2% for the US frontier, and 9.4% against 44.4% on ExploitGym. It came closest on ExploitBench, at 61.1% against 100%.

Is the Anthropic report independent?

No. Anthropic competes with Zhipu and sells defender access. CAISI's separate 17 September 2026 assessment reached the same four-month finding, which is why the shared conclusion is the safer one to rely on.

Sources

What each one is, and whose it is.

  1. OtherThe vendor’s own
  2. BenchmarkIndependent of the vendor
  3. Press reportIndependent of the vendor