Anthropic Warns About GLM-5.3 — Z.ai Responds With Its Own Defense Numbers
Anthropic published its Frontier Red Team evaluation of the Chinese open-weight model GLM-5.3 on September 29, 2026, and the conclusion is striking in two ways: the model has frontier-level cybersecurity capabilities according to Anthropic, and its safety filters can be bypassed with simple techniques in 64 to 100 percent of tests — or removed permanently for an estimated $1,200 in GPU time. Z.ai rejects the characterization and points to its own safety work, while NIST CAISI, according to available citations, has described GLM-5.3 as the most cyber-capable open-weight model to date. This is what the evaluation actually shows, what is independently verified — and what remains unresolved.
What Anthropic Published
Anthropic's post, signed by researchers including Robert Xiao and Tripp Gallagher, states that GLM-5.3 — developed by Zhipu AI, known outside China as Z.ai — has strong capabilities for autonomously building end-to-end cyber exploits, on par with Claude Mythos Preview. What sets GLM-5.3 apart from other frontier models, according to Anthropic, is that it was released without meaningful safeguards to restrict misuse (Anthropic; the phrasing here is reproduced in English translation).
The numbers behind the conclusion: On the ExploitBench benchmark suite, GLM-5.3 developed end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview managed the same in 56 of 410 — a nearly identical rate. Anthropic stresses that these tests were conducted in isolated, sandboxed environments, not against public systems on the internet.
It is also worth noting who is making the assessment: Anthropic is itself a commercial actor that competes with open-weight models, as BankInfoSecurity points out in its coverage. The capability figures should therefore be read as the company's own findings, not as neutral measurements.
The Bypass Ladder: 64, 92, 100 Percent
The most concrete part of the evaluation is an escalating series of bypasses of the model's safety filters, reported by TNW with reference to Anthropic:
- Telling the model it was an autonomous red-team agent got it to engage with the task 64 percent of the time.
- Pre-filling the model's reasoning so that it appeared to have already said yes raised the rate to 92 percent.
- A copy in which the refusal mechanism had been edited out of the weights engaged every time.
Anthropic writes that attackers can bypass GLM-5.3's safeguards between 64 and 100 percent of the time with simple techniques in their simulated tests, and that equivalent attacks failed against Anthropic's own, protected Claude models (Anthropic). BankInfoSecurity gives 63 to 100 percent — a minor discrepancy from Anthropic's primary source that remains unresolved but does not change the picture.
All tests were conducted in a sandbox, in other words. The findings do not demonstrate attacks against real systems on the internet, and neither Anthropic nor the coverage claims they do. But they show something else that matters for the open-weight debate: text-level filtering can be bypassed by changing the context, while filters embedded in the weights can be removed entirely when the weights are public.
Abliteration: The Filter Gone for ~$1,200
The most far-reaching claim in the evaluation concerns so-called abliteration — a technique in which the model's internal weight matrices are edited to remove its ability to refuse. According to SCMP via MSN, Anthropic spent roughly 2,200 GPU hours performing and testing abliteration on GLM-5.3. The result: the refusal rate fell from above 90 percent to between 2 and 12 percent on three safety benchmarks, with little capability loss.
TNW adds the economics: Anthropic's first attempt cost about $4,400, but the company estimates an experienced team would need around 600 GPU hours — roughly $1,200 — to do the same.
This is the point that changes the risk calculus for open-weight models. A vendor that only offers the model through an API can update filters, block misuse, and revoke access. When the weights are public, that lever no longer exists. GLM-5.3's safeguards function in practice as speed bumps: they slow down an inexperienced actor, but can be removed permanently with modest effort and equipment.
The Independent Corroboration — and Its Limits
The same month, NIST's Center for AI Standards and Innovation (CAISI) published its own assessment, dated September 17. CAISI found, according to Anthropic's summary, that GLM-5.3 is the most cyber-capable open-weight model released to date, and that it sits roughly four months behind the US frontier on an aggregate of CAISI's cybersecurity benchmarks (Anthropic).
This is the most important counterweight to the objection that Anthropic has a self-interest. But it is worth being precise: CAISI's assessment is cited here via Anthropic and secondary sources. CAISI's primary report should be read directly before drawing final conclusions about methods and uncertainties.
CAISI thus supports the capability assessment. It says nothing in itself about the quality of Z.ai's safeguards or how easily they can be removed — those are Anthropic's findings, conducted on the company's own terms.
Z.ai's Response
Z.ai has publicly rejected the characterization. Li Zixuan, the company's head of global operations, said according to SCMP that GLM-5.3 has already helped defend 389 open-source projects, with 4,249 potential vulnerabilities found so far (SCMP via MSN; reproduced in English translation).
These are the company's own figures, not independently verified. But they point to a real issue in the debate: the same capabilities that make a model dangerous in an attacker's hands make it valuable to defenders — who are, in practice, already using it.
There is also context on the release process. According to eWeek, Z.ai withheld the weights for two weeks after the August launch for additional testing and hardening. And when GLM-5.3 launched, the company claimed the model beat Anthropic's Mythos 5 on a central cybersecurity test — a claim from the same company now criticizing Anthropic's evaluation. The picture, in other words, is not that Z.ai ignored safety, but that the company and Anthropic assess what counts as sufficient very differently.
What It Means for the Open-Weight Debate
The juxtaposition of the three narratives — Anthropic's, CAISI's, and Z.ai's — illuminates the core dilemma in open-weight safety:
Weight-level filters are reversible when the weights are public. Anthropic's abliteration finding shows that even a model with safety training can have its refusal mechanism edited out for a few thousand dollars. That means "the model refuses" is not a durable property of an open model, but a property of one distribution of it.
Capability is approaching the frontier faster than safety practice. GLM-5.3 sits, according to CAISI, four months behind the US frontier. The gap between the most capable closed and open models is shrinking, even as open models lack the distribution control that makes API-based safety enforceable.
The conflict of interest runs both ways. Anthropic has an incentive to portray open-weight models as dangerous; Z.ai has an incentive to portray its model as safe and superior. CAISI's assessment is the closest thing to an independent reference point, but it covers capability, not the robustness of the safeguards.
What It Actually Costs to Do the Job for $20
A single example illustrates how concrete Anthropic's concern is. In one test case, the smaller GLM-5.3-Flash managed to chain together exploits for a known Chrome vulnerability, CVE-2026-11645, and another known flaw. It required 20 minutes of human attention and eight hours of model work, at an estimated API cost of $20.40 at Zhipu's prices — as reported by Anthropic (TNW).
This too happened in a sandbox against known vulnerabilities, not against real targets. But the arithmetic shows the mechanism: work an attacker previously had to spend days or weeks on can be partially automated, and the human's role shrinks to choosing targets and setting direction.
Open Questions
Several significant things remain unresolved:
- CAISI's primary report. The independent assessment is so far known via Anthropic's summary and secondary coverage. The methods and any caveats in the original report should be read before the conclusion is treated as final.
- Transfer outside simulation. All bypass and abliteration findings were demonstrated in isolated environments. How well the techniques work in practice against actual systems is not documented by the available sources.
- Z.ai's defense figures. 389 protected projects and 4,249 vulnerabilities found are the company's own claims without independent verification.
- Third-party confirmations. Some aggregator outlets have mentioned additional research-group confirmations beyond Anthropic and CAISI. AIMag has not seen documentation for these and therefore does not repeat them.
- 63 or 64 percent. BankInfoSecurity and Anthropic give slightly different lower bounds for the bypass rate. The discrepancy is small, but unresolved.
What does remain after the sources available is this: a capable open-weight model has been released with safeguards that one evaluating actor describes as easy to bypass and cheap to remove permanently. An independent government body has, according to available citations, reached the same capability assessment. The publisher does not dispute the capability — it disputes the interpretation. The debate that remains is not about how good GLM-5.3 is at finding vulnerabilities, but about who gets to decide.

