Back to Explore
NewsVibeFix 编辑部Updated Oct 10, 2026

After OpenAI's Agents, AWS Puts Zhipu's GLM-5.3 on the Bedrock Shelf: Chinese and American Models on the Same Cloud

Zhipu's flagship GLM 5.3 is generally available on Amazon Bedrock: a 753B-parameter MoE with a million-token context window and a leading 84.5 on the CyberGym security benchmark. AWS demoed it driving the open-source pentest agent Strix in an authorized security test. Behind the listing sits a revenue-share deal on invocation volume — Chinese and American models sold on the same cloud shelf, with Zhipu's Hong Kong shares jumping over 7% on the news.

Conceptual illustration of Zhipu GLM 5.3 joining the AWS Bedrock model shelf

One Shelf, Two Models: China and the US

In the first week of October, an unusual item appeared on AWS's model shelf. Earlier this month we covered "OpenAI Agents come to AWS" — the story of OpenAI's agents entering the AWS ecosystem. Days later, the other side of the Pacific took its turn: Zhipu (Z.ai)'s flagship model GLM 5.3 went generally available on Amazon Bedrock. An American closed-source agent and a Chinese open-weight model, listed on the same cloud shelf one after another, sold to the same enterprise customers. Different angles, same storyline: cloud vendors are turning the "model shelf" into the next competitive front.

A note on dating first — basic craft for news. The AWS blog post carries no inline date stamp. Three pieces of evidence pin the timing: the What's New announcement URL sits under /2026/10/; the blog's hero image filename contains 2026-10-05; and secondary corroboration lines up — Oossa labels it "Published October 5, 2026," and Z.ai's official X account shared it on October 6. Our call: around October 5. Every "October 5" below follows this dating.

AWS 官方博客配图:Bedrock 上 GLM-5.3 演示

What the GLM 5.3 Listing Brings: Specs and Integration

The hard numbers first, all from AWS's two official announcements. GLM 5.3 is Zhipu's flagship: a 753B-parameter mixture-of-experts architecture with roughly 40B parameters active per token, a 1-million-token context window, and up to 128K output tokens. Reasoning is on by default with selectable effort levels — you trade latency and token spend against task performance yourself.

Bedrock-side integration is thorough, the full "managed API" treatment: US and Global cross-Region inference profiles (us.zai.glm-5.3 / global.zai.glm-5.3) — you send requests to the source Region of your choice and Bedrock routes them securely; implicit prompt caching on by default, with explicit cache breakpoints available on the Responses / Chat Completions APIs; three service tiers — Flex (cheaper, for less time-sensitive work), Priority (latency-first, pricier), Standard (the default balance); and two API surfaces at once — OpenAI-compatible Responses / Chat Completions alongside Bedrock-native Invoke / Converse. AWS recommends the OpenAI-compatible set for new applications, where feature coverage is fuller.

The "reasoning always on" bit deserves its own mention: GLM 5.3 isn't a chat model that happens to reason — it's tuned as a reasoning model out of the box, with a per-task effort dial. Paired with a million-token context and 128K output, the positioning is clear: not for chit-chat, but for long jobs — reading whole repositories, running multi-round tool calls, emitting very long reports. The entire blog narrative orbits "long-horizon agentic tasks." Specs and story line up.

Two caveats, stated plainly. First, GLM 5.3 on Bedrock is currently limited to eligible enterprise customers — not click-and-go for everyone. Second, this is the GLM family's second Bedrock listing — GLM 5 arrived earlier this year; 5.3 is an upgrade of the same lineage. Zhipu reports a 50% gain over 5.2 on its internal coding benchmark and competitive showings on public coding benchmarks including DeepSWE, Terminal Bench 3.0, and FrontierSWE. Note the attribution: the AWS post cites these with a "Z.ai claims" qualifier, and we keep that qualifier — vendor claims and independent measurements are two different things.

The Most Interesting Detail: AWS Uses It as a "White-Hat Hacker"

A cloud vendor's model-launch post usually means spec tables, pricing, and a quickstart. What makes this one worth reading isn't the spec table — it's the demo AWS chose: GLM 5.3 driving a real agentic workflow — Strix, an open-source AI penetration-testing agent, running an authorized security test against OWASP Juice Shop, a deliberately vulnerable sample application.

Why security testing? Because Zhipu reported an eye-catching number: a leading 84.5 on the CyberGym security benchmark. AWS built the whole post's narrative around "security capability": Strix's official docs currently ship with GLM 5.3 as the default model, and pointing Strix at Bedrock moves inference inside the customer's own AWS account boundary — IAM permissions, audit logs, regional routing, the governance stack enterprises already have.

There's a reversal worth savoring here: agents like Strix used to call third-party inference services, with traffic leaving the building by design. Now AWS says: keep the traffic home, run it inside your own cloud account. The post also draws its red lines explicitly — test only applications you own or have explicit written permission to test; unauthorized testing is illegal in most jurisdictions and violates the AWS Acceptable Use Policy. A cloud vendor selling "offensive capability" has to lead with its compliance posture — that sentence is product design in itself.

One technical detail shows how deep the integration goes: Strix runs on LiteLLM under the hood, and at the time of writing LiteLLM couldn't resolve the bedrock/global.zai.glm-5.3 profile, so the post hands you a workaround — pin the Converse API route and the inference profile ARN directly. An official blog post walking you through a workaround by hand signals that AWS genuinely wants this agentic workflow to run, not just to display a logo.

The post also buries a product-line hint: if you'd rather not run Strix yourself, AWS offers Continuum, a managed service for on-demand penetration testing. The official line is "complementary" — open-source agents give you developer-driven, deeply customizable local testing; the managed service gives you assessments at scale. But the bigger game is visible: from "I sell you the model" to "I sell you model-powered services," with a toll collected at every layer. GLM 5.3 lands on the Bedrock shelf, gets adopted as Strix's default, then flows into Continuum's managed pipeline — a complete monetization chain taking shape.

Shelf Cadence: Two Chinese Labs in 48 Hours

Lay out the timeline and the cadence itself tells a story. Per The Inference: GLM-5.3 entered Bedrock on October 5, Kimi K3 followed on October 7 — flagship models from two Chinese labs, listed on the same American cloud shelf within 48 hours. That's not coincidence; it's Bedrock's shelf strategy. The number of models on the shelf is itself the moat: the more models listed, the stronger the case for going all-in on Bedrock, and the higher the switching cost.

More interesting still: American platforms now "sell Chinese intelligence" through two parallel channels. One is AWS-style cloud hosting — the model runs inside Bedrock's governance boundary, with revenue shared per invocation. The other is OpenAI's Codex reselling both GLM-5.3 and Kimi K3 through the inference provider Baseten — an American coding agent reselling Chinese weights. One sells "compliant inference," the other sells "a useful agent"; one's customer is the enterprise procurement department, the other's is the working developer. Both end at the same place: Chinese model intelligence, consumed on an American platform's invoice.

There's another calculation hidden in the details: the post casually mentions using Kimi K3 with OpenCode "as shown in our recent post," nudging readers toward the coding-assistant ecosystem. Listing the model is only step one — getting agent coding tools like OpenCode to call Chinese models on Bedrock by default is what actually drives invocation volume. And revenue share is paid on invocation volume: no calls, no split. Every one of these launch posts is, underneath, marketing for invocations.

"American Cloud Selling Chinese Intelligence": The Enterprise Math

Zoom out, and the procurement logic is blunt: enterprises no longer need to build and operate their own inference infrastructure to use a Chinese open-weight model. Billing goes through AWS, permissions through IAM, data boundaries and regional controls through Bedrock's existing compliance posture — "AWS security and compliance posture" is the single most important phrase in the What's New announcement. Customer data isn't shared with the model provider or used to train models; those promises live in Bedrock's documentation, not in Zhipu's sales deck.

This isn't the first case: DeepSeek-R1 became a fully managed Bedrock model back in March 2025. What makes GLM 5.3 different is twofold: it arrived with explicit commercial terms (next section), and the model itself is tuned for long-horizon agentic work — refactoring multi-file repositories, sustaining multi-hour agentic workflows without losing context, reasoning through complex systems problems with tool use at every step. The blog's opening names exactly these three workloads. In other words, Bedrock isn't selling a chat model; it's selling a workhorse that slots into enterprise agent pipelines.

"Governance" is a concrete thing here, not an adjective. The post's prerequisites section spells it out: calling GLM 5.3 requires the IAM permissions bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, and bedrock:CallWithBearerToken. Platform teams can use existing IAM policies to control precisely who calls, how much, and from which Regions; calls land in CloudTrail audits and Cost Explorer attribution. A Chinese open-weight model, slotted into the governance framework American enterprises already run — that's the technical footnote to "American cloud selling Chinese intelligence."

Revenue Share: Open Models' New Monetization Formula

What actually moved the industry here is the business model. Per The Inference's reporting: AWS put Zhipu's GLM-5.3 into Bedrock on October 5 and Moonshot's Kimi K3 on October 7 — both on revenue-share terms paid on invocation volume, not a one-time weight buyout or flat hosting fee. On the news, Zhipu's Hong Kong-listed shares (02513) jumped more than 7% (via Startup Fortune, citing Futu's news desk — secondary information).

Why does this model matter? Compare it with the classic open-weight trap: once weights are released, inference revenue flows to cloud vendors and inference providers while the lab itself captures nothing recurring. Revenue share reconnects the "selling intelligence" chain: the lab supplies weights and ongoing iteration, the cloud supplies infrastructure, customer trust, and the billing channel — the more invocations, the more both sides take home. Covering the Kimi K3 listing, one outlet put it well: Chinese open models going abroad are shifting from "maximizing deployment" to "monetizing at scale." The GLM 5.3 deal is another footnote to that sentence.

The split ratio, of course, is undisclosed. Back in late August there were reports that Moonshot had sought up to ~30% of K3-related service revenue from Microsoft, Amazon, and Google — but that was negotiation-stage reporting, never publicly confirmed by either side as an agreed figure. GLM-5.3's ratio is likewise undisclosed. Writing "unknown" plainly matters more than guessing a number.

The Inference added an even sharper observation: AWS listed GLM-5.3 and Kimi K3 two days apart, while OpenAI's Codex resells both through Baseten — "The tollbooth is American. The traffic is Chinese." Platforms keep the toll and the spread; the labs are interchangeable. Harsh, but it lays the shelf politics bare.

Capital markets' reaction is worth recording too. The same Inference piece gives valuation coordinates: Zhipu (listed in Hong Kong) at roughly $46.6 billion, Moonshot (private) at nearly $50 billion from its last round. Two Chinese labs earning per-call revenue on an American cloud while being priced by capital markets on "validated by American cloud" logic. A Bedrock listing announcement has, in a sense, become a credit endorsement for Chinese AI labs — unthinkable two years ago.

And a side note: that hero image in the AWS post — GLM 5.3 earnestly explaining symmetric vs. asymmetric encryption inside the Bedrock Playground — is itself a metaphor. A Chinese model teaching cryptography to American enterprises, inside an American cloud console. "Technology has no borders" is an old saying; "billing has borders, governance has borders" is this story's real subtext.

Bedrock 按调用量分成模式示意图

What It Means for Vibe Coders: A Middle Path

Back to our readers' perspective. Vibe coders running agentic workloads have long been stuck between two options: pay for closed APIs — expensive, with data leaving the building — or self-host open weights, owning all the hardware, ops, and debugging. GLM 5.3 on Bedrock offers a third path: open-weight economics plus cloud-vendor IAM, auditing, and regional controls — especially friendly to small teams, because compliance and governance come as a free ride rather than a build project.

The cost advantage isn't hand-waving. One independent benchmark (dev.to, secondary information) ran an instructive comparison: on a 660-file OWASP Java single-call code-review task, GLM 5.3, Kimi K3, and GPT-6 Sol tied within 1.5 points — yet GLM 5.3's input price was 56% of Kimi K3's and its output price 35%, making its cost per correct answer the lowest of the three ($0.0091 vs $0.0094). Opus 5.5 led outright by 7–9 points, at 3x+ the cost. The numbers say: in the "good enough" band, open-weight cost advantage is real.

Add Bedrock's prompt caching (agentic workflows resend large system prompts and repository context every turn; caching cuts latency and input cost directly) and three service tiers, and small teams get something new: the freedom to pick a price list by workload personality — Flex for unhurried nightly jobs, Priority for latency-sensitive interaction. That granularity never existed in the "one API key for everything" era.

The on-ramp is shorter than you'd think. The AWS post lays out three routes: easiest is the Bedrock console Playground — pick the model, send messages, zero code; for code, use OpenAI's Python SDK with aws-bedrock-token-generator minting short-lived tokens against the bedrock-runtime OpenAI-compatible endpoint — AWS explicitly recommends short-lived credentials over long-lived API keys; and if you already live in an agent coding tool like OpenCode, point it at Bedrock as a model source. Explicit prompt caching needs prompt_cache_options declared on the request with breakpoints — each breakpoint needs at least 1,024 tokens to qualify — the single most worthwhile optimization to tune for long-context agent work.

One gate to flag: GLM 5.3 is currently open only to eligible enterprise customers. In other words, this isn't consumer acquisition — it's directed enterprise supply. Enterprise buyers pass eligibility checks; the lab gets high-quality, high-willingness-to-pay invocation volume. Which also explains why the revenue-share math works: the split isn't over retail pennies, it's over the AI-inference line of enterprise budgets.

Variables and Risks: Washington Is Watching

None of this is risk-free. Multiple outlets covering the Kimi K3 listing flagged the same variable: Washington's scrutiny of Chinese AI models on data security, IP, and national security grounds is the biggest uncertainty over how far these partnerships can expand. Revenue share negotiates the commercial terms; policy may set the ceiling.

For enterprise buyers that means two things. First, contract continuity belongs in the evaluation: a model listable today could be delisted tomorrow on policy grounds. Bedrock's cross-Region inference and regional controls solve data-compliance questions, but not the "can this model still be sold" question. Second, keep an abstraction layer in your architecture — and here's where OpenAI-compatible APIs pay off: today you call global.zai.glm-5.3, tomorrow you swap models by changing a model ID, not rewriting your call stack. Treating models as replaceable parts rather than articles of faith is the only posture that survives policy cycles.

My read: over the next year, "where the model runs" will matter as much as "which model." Open weights solved "can you use it"; revenue share solves "can the lab keep making money"; managed hosting like Bedrock solves "does the enterprise dare use it." Put the three together and the middle path for agentic workloads finally walks. AWS just put two of the three puzzle pieces — hosting and revenue-sharing — on the table in one move.

And back to the opening image: American closed-source agents and Chinese open-weight models, sold side by side on the same shelf. A decade ago cloud vendors competed on VMs and storage; five years ago on data lakes and serverless; now the contest is "whose shelf holds the most intelligence." AWS listed GLM-5.3 and Kimi K3 within days of each other and bound Chinese labs to its own billing via revenue share — this isn't a technology story, it's a channel statement. The future of AI competition will be, to a large extent, a competition of channels. Whoever owns the enterprise procurement doorway owns the pricing power. Models will keep becoming interchangeable. Shelves won't.

Sources and Dating Note

Per house standards: primary sources are the AWS Machine Learning Blog post "Introducing GLM 5.3 on Amazon Bedrock" and the AWS What's New announcement; secondary information (listing timeline, revenue-share terms, Zhipu's >7% share jump, Kimi K3 timing, Baseten reselling, dev.to benchmark figures) is attributed inline throughout. Dating follows the first section: judged to be around October 5 on combined evidence, stated transparently.

Sources

Browse projectsPublish your project

Related articles

Illustration of 2,000 AI agents collaboratively rewriting the Prime Agent codebase from TypeScript to Rust
News
2,000 Agents Rewrote Themselves in Rust: Prime Intellect's Two-Week Dogfooding Experiment

Prime Intellect had Prime Agent orchestrate 2,000+ agents to rewrite itself from TypeScript into Rust in two weeks, burning 200B+ tokens across 10,000+ sandboxes. The real story is not the 14x speedup but the honest methodology: a root agent that writes no code, verification separated from implementation, and an open admission that passing scripted parity tests does not mean production-ready.

AI CodingOpen-source ProjectsIndustry Trends
Illustration of parallel AI agent tasks scanning a map service, with public scan reports forming a visible trail
News
Fleet, Not Swarm: A Chinese Agent Fleet Surfaces on Tencent Cloud After 2,048 Amap Scans

Between Sept 28 and Oct 4, independent researchers at Swarmchasers found 2,048 public urlquery.net scan reports targeting Alibaba's Amap, peaking at 1,810 in a single day. The agents ran on Tencent Cloud behind a proxy named hysandbox-ats, labeled themselves 'claude' while code fingerprints pointed to Hunyuan and GLM models, and kept no coordination channel at all. This is why the researchers insist on 'fleet,' not 'swarm' — and why agent observability cuts both ways.

Security & PrivacyIndustry TrendsAI Coding
Kimi K3 logo beside the OpenAI Codex developer interface and a billing invoice graphic
News
Kimi K3 Enters OpenAI's Enterprise Codex Channel — the First Chinese Open Model Inside OpenAI Billing

Moonshot AI's Kimi K3 has entered OpenAI's enterprise Codex channel via US inference provider Baseten. Enterprise customers can now burn existing OpenAI spending commitments on the Chinese open model — no new supplier contract needed. Sina Finance calls it the first Chinese open model to enter OpenAI's enterprise billing system. This piece unpacks the three-way split (Baseten/Codex/OpenAI billing), the Bedrock revenue-sharing lead-up, and what billing-neutral model choice means for developers.

Model UpdatesIndustry TrendsAI Coding