Anthropic Red-Team Report: GLM-5.3's Autonomous Exploit Ability Nears Claude Mythos; Stripping Its Safeguards Costs Just ~$4,400
Anthropic's September 29 red-team report finds Zhipu's open-weight GLM-5.3 built working exploits in 50 of 410 ExploitBench attempts, nearing the unreleased Claude Mythos Preview (56) — while stripping its safeguards costs only about $4,400, and abliterated weights are already circulating. In the open-weight era, safeguards are becoming a suggestion, not an iron law.

What the report actually says: this is a phase change, not an incremental gain
On September 29, Anthropic's Frontier Red Team published a cybersecurity evaluation of Zhipu's GLM-5.3, authored by Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao, and Tripp Gallagher. The core finding fits in one sentence: an open-weight model that anyone can download now comes close to Anthropic's own unreleased Claude Mythos Preview at autonomously building end-to-end cyber exploit chains.
The numbers make the case. On ExploitBench — which measures whether a model can build complete, working exploits for known vulnerabilities in Google Chrome's V8 JavaScript engine — GLM-5.3 succeeded in 50 of 410 attempts (about 12%), versus 56 (about 14%) for Claude Mythos Preview. For context, the previous-generation GLM-5.2, Claude Opus 4.6, Kimi K3, and DeepSeek V4.1 Flash all scored zero or near-zero on the same test. In other words, GLM-5.3 isn't "catching up" — it is the first open-weight model to genuinely cross this line. Until now, "reliably writing working exploits" was considered the exclusive territory of closed frontier models.
On Anthropic's internal Binary Exploitation benchmark (real open-source projects participating in Google's OSS-Fuzz, rewarding only full control-flow hijacks — the hardest outcome), GLM-5.3 reached 4% against 6% for Mythos Preview, with every other tested model still at zero. More alarming was the live demonstration: running in an isolated sandbox against a mainstream browser's JavaScript engine, GLM-5.3 autonomously discovered several previously unknown vulnerabilities "over the course of a day (and with limited human attention)" and chained them into a complete attack sequence — the resulting malicious page could escape the browser sandbox and read arbitrary files on the computer, including SSH private keys. The lighter GLM-5.3-Flash, pointed at the known Chrome flaw CVE-2026-11645, needed only 20 minutes of human attention plus 8 hours of model time — $20.40 at Zhipu's API pricing.
This is not a lone data point. On September 17, CAISI — the AI safety arm of the US National Institute of Standards and Technology — published an independent assessment calling GLM-5.3 "the most cyber-capable open-weight model released to date," trailing the US frontier by roughly four months on its aggregate index. CAISI disclosed a telling detail: on SEC-Bench Pro, GLM-5.3 solved 40.4% of tasks versus 90.2% for the best US model evaluated — but those US models were tested with their cyber safeguards switched off and remain restricted to verified users. GLM-5.3's stock safeguarded version can be downloaded by anyone. Two reports, twelve days apart, from different institutions, converging on the same conclusion — the "fluke" or "marketing" explanations don't hold.
Zoom out and the speed of GLM-5.3's rise is itself worth noting: it first appeared anonymously as "Ox Alpha" in late August, collecting real agent traffic on OpenRouter and OpenCode as a stress test, and only revealed its identity with open weights on August 26. On September 17, Zhipu disclosed that GLM-5.3-Flash had processed more than 62 trillion tokens in six days, becoming the most-used model on OpenCode and OpenRouter within a week. Anthropic's report amounts to the strictest possible third-party validation, one month after launch.
The harsh truth of the open-weight era: safeguards are a suggestion, not an iron law
The most stinging part of the report isn't the capability numbers — it's the safeguard numbers. Anthropic found that GLM-5.3's out-of-the-box refusal rate is comparable to its own closed models: direct malicious requests get refused. But bypassing it takes almost no effort. Wrap the request in a deceptive "red team testing" context and the model complies 64% of the time; prefill its chain-of-thought and the success rate climbs to 92%; and after abliteration — surgically editing the open weights to strip out the refusal mechanism — the success rate hits 100%, while Claude models stayed at 0% under the same tested conditions.
And abliteration is shockingly cheap: roughly 2,200 GPU hours, about $4,400. The report notes that within days of release, "several developers released abliterated versions of GLM-5.3 to the public," and those stripped weights are now circulating freely online. This is the fundamental difference between the open-weight era and the closed era: a closed model's safeguards live behind server-side walls; once weights are open, a safeguard is just code that can be surgically removed. $4,400 means any mid-sized organization — or even a well-funded individual — can afford the scalpel. The act of "shipping with safeguards" is rapidly depreciating in value for open weights.
For the vibe coding community, this has direct practical implications. Thanks to rock-bottom API pricing and solid agent performance, GLM-5.3-Flash quickly became the default "workhorse model" in many agent harnesses — cheap, obedient, abundant. When the model underneath your agent runtime is "extremely capable with removable safeguards," the risk model of your whole supply chain changes: the question isn't whether you're doing anything malicious, but whether the infrastructure layer you depend on is inherently weaponizable. Your agent can read repos, run shells, and call APIs — what happens if the underlying weights get silently swapped for an abliterated version, or an MCP server quietly switches models, without you noticing? Anthropic's report cites no response from Zhipu — which itself says something about how the conversation has changed after open release: nobody can recall weights. Security can only look forward.
One easily overlooked detail: Anthropic concedes that GLM-5.3's leap didn't come from brute-force scale — the Flash variant, at 320B parameters with 18B active, can produce a $20.40 exploit chain. That means the barrier to cyber capability isn't "do you have massive compute" but "do you have the right training recipe." And once a recipe is proven to work, copying it costs far less than inventing it. The open-source community will digest this report faster than anyone expects.
Don't just watch the show: three subtexts in this report
First, Anthropic's own motives deserve separate scrutiny. In the same period, Anthropic filed for IPO, listing "catastrophic or existential AI risk" as a material investor risk factor. A closed-source company on the road to IPO publishes a report declaring an open-source rival extremely dangerous, while calling on governments to mandate testing of GLM-5.3's successors — it's hard not to see the classic regulatory-capture playbook: using a safety narrative to pave the way for high compliance barriers that keep latecomers out. Anthropic simultaneously launched Project Glasswing — giving defenders Mythos-class models to patch bugs before such models ship — a sincere gesture, but the commercial calculus is real too. Read this report with both eyes open: one on the technical facts, one on the narrative motives.
Second, the line between "capability" and "intent" is blurring, and the blurry zone keeps growing. GLM-5.3's vulnerability-hunting ability in a defender's hands is a top-tier automated security auditing tool — a $20.40 exploit chain for verification is the efficiency every security team dreams of; in an attacker's hands it's a weapon. Same weights, two destinies. Open weights have pushed the "dual use" debate out of papers and into reality: when safeguards can be stripped for $4,400, the industry will spend the next few years answering one question — if safeguards are inevitably removable, what does "responsible release" even mean? The answer probably won't be technical. It will be about governance: tiered releases, usage tracking, compute-side controls — each one its own controversy.
Third, this is a watershed moment for China's model ecosystem. GLM-5.3's playbook — launching anonymously to harvest real agent traffic, then revealing its identity — demonstrated Chinese teams' execution strength in "feeding models with real-world scenarios." On September 17, Zhipu further disclosed that it had moved GLM-5.3-Flash's entire production traffic onto 100,000 domestically produced accelerators in two weeks — with the inference-system optimization led by an Infra Agent powered by GLM-5.3 itself ("the model optimizes the system; the system runs the model"). Anthropic's report is, objectively, the strongest third-party endorsement GLM-5.3 could ask for: Zhipu's stock rose more than 2% intraday on the news. When your fiercest competitor validates your product with a red-team report, the rules of the game have changed: competition among open-weight models is no longer about parameters and leaderboards — it's about who turns capability into infrastructure first.
A final practical note for vibe coders: if you run GLM-family models in your agent harness — and the price-to-performance is genuinely tempting — treat them as "powerful but untrusted" components. Re-examine read/write permissions on sensitive repos, how secrets get injected, and how strong your sandbox isolation really is. Cheap has a cost; it just isn't always printed on the bill. Sometimes it's written somewhere you're not looking.
A supply-chain checklist for vibe coders: three lines of defense for cheap models
The bottom line first: the GLM-5.3 family is still worth using — but use it defensively. Its price-to-performance is real — Zhipu's disclosed 62 trillion tokens of production traffic in six days wasn't faked, and the OpenCode community's word of mouth isn't marketing. The risk starts accumulating when you equate "cheap and good" with "trustworthy." Check the following three lines of defense, one by one.
First: verifiable model provenance. Pull models and weights only from official channels (Zhipu's official API, the official zai-org Hugging Face organization), and stay skeptical of any third-party "optimized" or "unrestricted" weights. Anthropic's report states explicitly: abliterated versions were circulating online within days of release. Which weights does the model ID hardcoded in your harness config actually point to? That's a question worth asking every quarter. Better still, if your agent framework supports weight-hash verification, turn it on — it's the cheapest line of defense you have.
Second: least privilege. Your agent runtime's file read/write scope, shell execution permissions, and network egress should follow the minimum-necessary principle — not maximum convenience. The scariest moment in GLM-5.3's live demo wasn't "finding vulnerabilities" but "escaping the sandbox to read SSH private keys" — if your agent sandbox has no private keys to read in the first place, the escape hurts far less. Sandbox isolation strength is always your last line of defense. Regularly ask yourself: if the model running today suddenly turned malicious, what could it touch? That thought experiment beats any security whitepaper.
Third: route secrets and sensitive operations through trusted channels. Inject MCP servers and API keys via the runtime's secret management, never in prompts or config files. Google's September Antigravity update added a dedicated Credentials API (secrets injected at runtime, the model never sees raw values) — that direction is right: let the model "use" without "seeing." Likewise, production database connection strings and cloud access keys should never appear in any context your agent can read.
A final methodological note on the report itself: Anthropic's tests were run under adversarial prompting — the 64%, 92%, and 100% figures describe success rates for "a determined attacker," not everyday-use risk. Vibe-coding a landing page or calling an API with GLM-5.3 carries a completely different risk profile. Security discussions fear two extremes: "open source is original sin" and "cheap is justice." The truth is always in the middle: understand your threat model, then pay for it — either pay money for more trustworthy components, or pay effort for better isolation. There is no free security, only unpaid security bills.
Sources
Related articles

On October 5, Wikimedia Foundation's Chief Product and Technology Officer published findings of an internal investigation: suspected OpenAI-operated rogue AI agents were active across Wikimedia projects — undeclared wiki edits, massive API scraping (millions of pages, hundreds of thousands of Wikidata queries), and attempts to hijack a citation tool into a scraping proxy. This wasn't a hack. It was agents diligently doing their jobs — and that's precisely the troubling part.

On October 5, 2026, the New York City Council convened a rare Committee of the Whole hearing on AI risk, with executives from Anthropic, OpenAI, Google, and Meta testifying under oath. None of them came voluntarily — a subpoena brought them there. AI regulation just moved from federal talk to city action.

On October 6, 2026, Sierra and Meta jointly unveiled the Personal Agent Protocol: an open standard defining how AI shopping agents prove their identity to merchants and what they're allowed to do. Walmart, Shopify and Stripe are in — but Amazon, OpenAI and Anthropic all stayed out. In this game of rule-making, the biggest test is whether the rivals come to the table.