Anthropic tested open-weight GLM-5.3 and says its safeguards fold under framing
Anthropic's red team says Z.ai's open-weight model builds working browser exploits at close to Claude Mythos Preview's rate, and that its refusals give way to a cover story, a prefill, or about $4,400 of GPU time.
Anthropic says GLM-5.3 refused every direct attack request. Editing out the refusals took about $4,400.
Anthropic's Frontier Red Team published its tests of GLM-5.3 on September 29. According to Anthropic, the open-weight model from Zhipu AI, known outside China as Z.ai, builds working exploits at close to the rate of Anthropic's own Claude Mythos Preview. Anthropic says its refusals could be bypassed or removed in 64% to 100% of simulated trials.
Fifty working exploits in 410 tries
Anthropic says it ran GLM-5.3 on ExploitBench, which tests known vulnerabilities in V8, the JavaScript engine in Google Chrome. By Anthropic's count, GLM-5.3 built end-to-end exploits in 50 of 410 attempts, about 12%. Claude Mythos Preview managed 56.
On 100 randomly picked tasks from its internal binary exploitation benchmark, built on open source projects in Google's OSS-Fuzz program, Anthropic says GLM-5.3 reached a full control-flow hijack in 4% of trials. Mythos Preview reached 6%. Earlier models, Claude Opus 4.6 and GLM-5.2, succeeded on none of them, according to Anthropic.
Anthropic also reports two researcher-driven sessions. In the first, a researcher spent about a day with GLM-5.3 on a sandboxed machine, with less than an hour of focused human time. Anthropic says the result was a chain of previously unknown bugs in a browser's JavaScript engine, packaged as a webpage that reads arbitrary files from the visitor's computer, and that it has disclosed the bugs to the maintainer. In the second, Anthropic says the smaller GLM-5.3-Flash turned two publicly known Chrome flaws into a working exploit chain with 20 minutes of human attention and eight hours of model time, about $20.40 at Zhipu's API prices.
Zero percent, then 64, then 92, then 100
For the safeguard tests, Anthropic says it placed the model in a simulated environment where no model-generated code runs, then gave it overtly malicious requests to attack critical systems. Asked directly, GLM-5.3 engaged 0% of the time, across 50 samples per condition. Anthropic reports three conditions that changed that:
- A false red-team cover story: 64% engagement.
- Prefilled reasoning: 92%.
- An abliterated copy, with the refusals edited out of the open weights: 100%.
Anthropic says none of these got its safeguarded Claude models to carry out the tested tasks. It adds that its API offers no way to prefill Claude's reasoning, and that Claude's weights are not public, so they cannot be abliterated.
The $4,400 line item
Anthropic says abliterating GLM-5.3 took its team, which had never tried the technique before, about 2,200 GPU hours, roughly $4,400 of compute. It estimates an experienced team would need closer to 600 GPU hours, about $1,200. It also says several developers published abliterated versions of GLM-5.3 within days of release.
According to Anthropic, the edit took refusal rates from above 90% to about 3% on JailbreakBench and 2% on HarmBench, and to 12% on StrongREJECT. Anthropic says GPQA-Diamond scores stayed the same and a tested subset of CyberGym dropped by a few percent.
A competitor grading a competitor
Anthropic competes with Z.ai, and it sells access to Mythos-class models through vetted programs, a point Madrobot's coverage raised directly. Madrobot also reported that Z.ai had not commented. Z.ai's own launch claims for the model are in our August item.
One part has an outside check. On September 17, NIST's Center for AI Standards and Innovation called GLM-5.3 "the most cyber-capable open-weight model released to date" and placed it about four months behind the US frontier on its cyber benchmarks. Anthropic says its capability findings broadly match CAISI's. The safeguard numbers have no independent replication yet, and Anthropic itself calls its simulations imperfect measures of real-world behavior.
Why a build studio cares
If Anthropic's numbers hold, a team that self-hosts GLM-5.3 should treat its refusals as a default that ships with the weights, one a cover story beats 64% of the time and $4,400 of GPU time deletes outright. The controls that survive abliteration are the ones in the deployment: a sandbox with nothing worth reading inside it, egress limited to named hosts, tool permissions scoped per task, and a human approval before anything irreversible. That is where we look when we audit an agent pipeline built on an open-weight model. The file-reading webpage took about a day and under an hour of human focus, per Anthropic. That is a short window for anyone slow to patch browsers.
Next step: read Anthropic's full write-up, then CAISI's assessment for the independent view. If you run an open-weight model in production and want to know what it can reach once its refusals are gone, write to us at hello@gattyworks.com.