ResearchResearch 5 min read

Anthropic Warns GLM-5.3 Spreads Advanced Cyber Exploits

Anthropic’s 29 September 2026 research post on Zhipu AI (known outside China as Z.ai) and its open-weight GLM-5.3 argues that the model nearly matches Claude Mythos Preview at building cyber exploits while shipping without meaningful misuse safeguards.

PC

PromptCrates Editorial

Staff Writer

0 0
Anthropic Warns GLM-5.3 Spreads Advanced Cyber Exploits

Anthropic on 29 September 2026 published a research post on Zhipu AI (known outside China as Z.ai) and its open-weight GLM-5.3, arguing that the model now matches a critical cyber-exploit threshold previously associated with Claude Mythos Preview—while shipping without meaningful misuse safeguards. Authors Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, and Tripp Gallagher write that attackers can bypass GLM-5.3’s safeguards between 64 percent and 100 percent of the time with simple techniques in Anthropic’s simulated tests, unlike safeguarded Claude models in the same harness.

Capability findings on open weight cyber tools

Five months earlier, Anthropic said Claude Mythos Preview was the first model that could autonomously build sophisticated end-to-end cyber exploits, and it limited release through Project Glasswing so trusted defenders could find more than 10,000 vulnerabilities in critical software before peers caught up. The new post says those peer models have arrived. On ExploitBench, which measures end-to-end exploits against known V8 vulnerabilities in Google Chrome, GLM-5.3 succeeded in 50 of 410 attempts versus 56 of 410 for Claude Mythos Preview. On Anthropic’s internal Binary Exploitation benchmark of 100 OSS-Fuzz-linked tasks, GLM-5.3 achieved full control-flow hijacks in 4 percent of trials versus 6 percent for Mythos Preview, while earlier Claude Opus 4.6 and GLM-5.2 scored zero.

Human-in-the-loop sessions sharpened the warning. In a sandboxed Linux browser build, GLM-5.3 found previously unknown JavaScript-engine flaws and chained them into a webpage exploit that read arbitrary files from the visitor’s machine; Anthropic says it disclosed those issues to the maintainer. In a second session, GLM-5.3-Flash turned public details of Chrome CVE-2026-11645 plus another known flaw into a reliable ARM64 exploit chain that bypassed pointer-authentication hardening, using about 20 minutes of human attention, eight hours of model work, and roughly $20.40 at Zhipu API prices.

NIST’s Center for AI Standards and Innovation (CAISI) published its own GLM-5.3 assessment on 17 September, calling it the most cyber-capable open-weight model released to date and placing it about four months behind the U.S. frontier on CAISI’s aggregate cyber benchmarks. Anthropic says its capability findings broadly match CAISI’s, with the crucial distribution difference that anyone can download GLM-5.3 while many U.S. frontier cyber evaluations use safeguarded or vetted-only configurations attackers cannot casually reach.

Safeguard bypasses and defender implications

Out of the box, GLM-5.3 often refuses clearly harmful requests, but Anthropic reports the protections are thin. Abliteration—a refusal-reduction edit possible because weights are public—dropped refusal rates from above 90 percent to about 3 percent and 2 percent on JailbreakBench and HarmBench and to 12 percent on StrongREJECT, at an estimated $4,400 compute cost for Anthropic’s inexperienced-with-the-technique team (closer to $1,200 for experts). Capability scores stayed largely intact. Without abliteration, a deceptive red-team cover story engaged the model 64 percent of the time, prefilling thinking tokens reached 92 percent, and abliterated weights reached 100 percent. None of those techniques got safeguarded Claude models to carry out the harmful remote-attack tasks in the same simulated setup.

Anthropic argues this is a meaningful step change in freely accessible attacker capability, while also urging defenders to use the best tools that meet their needs and expanding trusted access to Claude cyber capabilities. The company points to Project Glasswing and Patch the Planet as partial head starts that remain incomplete. Governments, it says, should safety-test sufficiently capable models, including GLM-5.3 successors. Readers tracking Anthropic’s wider risk posture can juxtapose this post with Anthropic’s IPO prospectus risk language and independent intelligence-explosion oversight research.

What security teams should do this week

Inventory whether any internal red-team, research, or contractor workflows already pull GLM-5.3 or abliterated variants from public hosts, and treat those endpoints as high-risk tools even if marketed for coding help. Second, accelerate patching and monitoring for browser engines, drivers, and network-facing services of the class Anthropic says the model reached in short sessions—without waiting for every CVE to trend. Third, revisit open-weight allow-lists: “open for transparency” is not the same as “safe for unvetted agents on corporate laptops.”

Policy leads should also update board briefings. CAISI’s four-month lag figure and Anthropic’s bypass rates are concrete enough for non-technical directors, and they pair with legislative debates covered in PromptCrates’ Human Control Over AI Act reporting. Finally, if your defenders need frontier cyber assistance, evaluate vetted access programs rather than assuming the open-weight shortcut is equivalent once abliteration exists.

Vendors and open-source hosts share responsibility in the distribution path. When abliterated weights appear within days of a release, mirror operators and internal artifact registries become part of the security perimeter. Require signed provenance for any GLM-family checkpoint used in corporate environments, and block anonymous downloads of abliterated builds on developer laptops the same way you would block cracked commercial compilers. Document exceptions for vetted research under isolated networks with no egress to production identity systems.

Incident-response playbooks need a specific branch for AI-assisted exploit development. If a browser or driver zero-day arrives with evidence of heavy model-assisted iteration, assume more variants will follow quickly and shorten the window between disclosure and mass scanning. Coordinate with bug-bounty and vendor PSIRT contacts early, citing the public Anthropic and CAISI assessments so urgency is not dismissed as vendor marketing.

Primary reporting for this article: Anthropic’s 29 September research post on GLM-5.3 and the spread of advanced cyber capabilities, including ExploitBench and Binary Exploitation figures, human-session narratives, CAISI’s 17 September assessment, abliteration cost and refusal metrics, and the 64 percent, 92 percent, and 100 percent bypass conditions.

researchAnthropicGLM-5.3cybersecurity

Related articles