Anyone Can Download What Attackers Could Not Build — Anthropic's GLM-5.3 Cyber Assessment
Five months ago Anthropic announced Claude Mythos Preview, a model that could autonomously build sophisticated end-to-end cyber exploits, and released it only to vetted defenders through Project Glasswing. On 29 September 2026 its Frontier Red Team published "GLM-5.3 and the spread of advanced cyber capabilities" with an uncomfortable update: the capability has proliferated, exactly as predicted — into GLM-5.3, an open-weight model from Zhipu AI (Z.ai) that anyone can download, and whose safeguards fall to trivial bypasses at rates of 64% to 100%.
This is not a warning about a hypothetical. The report puts numbers on exploit development, safeguard removal, and dollar cost, and its capability findings broadly match an independent assessment NIST's Center for AI Standards and Innovation (CAISI) published on 17 September. Two separate labs, one conclusion: the most cyber-capable open-weight model released to date is now freely available.
Capability: at parity with the restricted frontier
On ExploitBench, which measures exploitation of known vulnerabilities in the V8 engine used by Google Chrome, GLM-5.3 developed working end-to-end exploits in 50 of 410 attempts — against 56 of 410 for Claude Mythos Preview. On Anthropic's internal Binary Exploitation benchmark (100 randomly selected tasks from Google's OSS-Fuzz projects, full credit for a complete control-flow hijack), GLM-5.3 succeeded in 4% of trials versus 6% for Mythos Preview. The models it left at zero are the telling part: Claude Opus 4.6 and GLM-5.2 succeeded in none of them. A threshold has been crossed between generations, not within them.
The human-in-the-loop results are sharper. A researcher paired with GLM-5.3, working with less than an hour of total human focus, found previously unknown vulnerabilities in a browser JavaScript engine and chained them into a working exploit — a webpage that reads arbitrary files off a visitor's machine, demonstrated stealing an SSH private key. Anthropic says it disclosed those vulnerabilities to the maintainer, and that the same sessions surfaced exploitable flaws in wireless and graphics drivers and network-facing device software, now under review for disclosure. The model also appears alongside Moonshot's Kimi K3 and DeepSeek's V4.1-Flash in the report's scaling charts — this site has tracked both lines of evaluation through the UK AISI work on Kimi K3 and the DeepSeek Hermes campaign.
The $20.40 number that prices the N-day pipeline
The report's most quotable experiment used the smaller GLM-5.3-Flash against CVE-2026-11645, a recently disclosed Chrome flaw. Given public details of that CVE plus one more known flaw, and no significant direction, the model chained the two into a reliable exploit for an ARM64 target that bypasses pointer-authentication (PAC) hardening. Cost: 20 minutes of human attention, eight hours of model time, $20.40 at Zhipu's API prices.
That figure belongs next to the finding in the Artificial Analysis Cyber Index, published the day before: models are getting fast at the find-and-fix loop while patch deployment still lags. A $20 N-day-to-exploit pipeline compresses the attacker's side of that gap to hours. Every day a patch sits undeployed is now a day a $20 run can weaponise.
Safeguards: three bypasses, all simple, one total
GLM-5.3 ships with refusal-style safeguards — ask for something clearly harmful and it will often refuse. Anthropic tested three ways past them in a simulated environment with overtly malicious requests:
- A deceptive prompt — telling the model it is an autonomous red-team agent on an exercise — engaged it 64% of the time.
- Prefilling its thinking tokens so it appears to have already decided to proceed — 92%.
- Abliteration, the standard refusal-removal edit, possible because the weights are open — 100%.
None of these worked against safeguarded Claude models in Anthropic's testing: deceptive prompts were blocked, the API offers no thinking-prefill path, and closed weights cannot be abliterated. The abliteration economics are the grimmest detail. Anthropic's team, which had never attempted the technique, did it in about 2,200 GPU hours (~$4,400) and estimates an experienced team needs ~600 hours (~$1,200). Refusal rates collapsed from above 90% to 3% on JailbreakBench, 2% on HarmBench, 12% on StrongREJECT — while GPQA-Diamond scores stayed identical, meaning capability survived the edit intact. Several developers published abliterated copies within days of release.
What to do
- Price N-day exposure in hours, not weeks. If your patch cycle for browser and edge components is measured in weeks, the $20.40 experiment says the weaponisation side now runs in under a day. Prioritise Chrome/V8 and internet-facing patch queues first.
- Treat open-weight capability as the attacker's baseline. CAISI's assessment — most cyber-capable open-weight model to date, roughly four months behind a US frontier attackers cannot readily access — means threat models assuming "attackers lack exploit-development capability" need rewriting.
- Give defenders the same tools. Anthropic's own conclusion: defenders should work with frontier models at least as capable as their adversaries'. Project Glasswing's 10,000+ vulnerabilities found by vetted defenders is the template; the report explicitly urges expanding trusted access, including Mythos 5.1-class models, to more entities.
- Demand independent evaluations. Without CAISI-quality third-party testing, the report argues, developers will not see the impact of these capabilities until too late. Procurement and policy teams should ask vendors for benchmark-backed cyber-capability and safeguard data, not refusal-rate marketing.
- Watch the disclosure pipeline. Anthropic's still-embargoed driver and device-software findings will become patches and then N-days. Track the Frontier Red Team feed the way you track CISA KEV — the second half of this story arrives as CVEs.
Our verification was documentary and primary-source-led. We fetched Anthropic's full report page and extracted the ExploitBench (50/410 vs 56/410), Binary Exploitation (4% vs 6%, predecessors at zero), CVE-2026-11645 experiment ($20.40, 20 minutes human, 8 hours model, ARM64 PAC bypass), abliteration economics (2,200 GPU hours, refusal-rate collapses, GPQA-Diamond parity), the three bypass rates (64/92/100%), the CAISI 17 September findings, and the Project Glasswing figures directly from the published text. Secondary coverage was used only to confirm publication date and authorship. We did not run any model, reproduce any bypass, or develop exploit code.
Sources:
- Anthropic Frontier Red Team — "GLM-5.3 and the spread of advanced cyber capabilities" (29 September 2026; ExploitBench, Binary Exploitation, human-in-the-loop zero-days, CVE-2026-11645 experiment, safeguard bypasses, abliteration economics, CAISI comparison, Glasswing)
- eSecurity Planet — Anthropic finds GLM-5.3 nears Mythos Preview in exploit tests (independent summary of benchmark figures)
- Resilient Cyber Newsletter #116 — GLM-5.3-Flash CVE-2026-11645 chain corroboration ($20.40 figure)