2,000 Agents Rewrote Themselves in Rust: Prime Intellect's Two-Week Dogfooding Experiment
Prime Intellect had Prime Agent orchestrate 2,000+ agents to rewrite itself from TypeScript into Rust in two weeks, burning 200B+ tokens across 10,000+ sandboxes. The real story is not the 14x speedup but the honest methodology: a root agent that writes no code, verification separated from implementation, and an open admission that passing scripted parity tests does not mean production-ready.

Prime Intellect has published Rewriting Prime Agent in Rust on its official blog, announcing that its open-source coding agent, Prime Agent, has been rewritten from TypeScript into Rust from the ground up — and the rewrite was executed by Prime Agent itself. A note on dating: the blog post itself carries no dateline. The October 9, 2026 date is corroborated by RuntimeWire's original reporting (stamped "Published Oct 9th, 2026, 7:54pm CT") and a same-day post from Prime Intellect's official X account.
This deserves its own piece not because "another project switched to Rust," but because it is the largest publicly documented "agents writing agents" experiment to date. Two weeks, more than 2,000 agents, 10,000+ Prime Sandboxes, and 200 billion+ tokens burned on Prime Inference's GLM-5.3 endpoint — the numbers alone read like a paper-grade engineering experiment log. What is rarer is that Prime Intellect wrote up the methodology, the failures, and the blind spots. That honesty is worth more than the 14× performance figure.
The bill first: every number matters
Prime Agent launched in August; by the time of the post it had been downloaded more than 300,000 times and processed over 8 trillion tokens. The rewrite's resource consumption is broken into two phases, with numbers precise to two decimal places:
- Rust implementation phase: 192.99B tokens, 1,981 agents;
- Performance hillclimbing phase: 35.70B tokens, 228 agents.
All told: 2,000+ agents, 10,000+ Prime Sandboxes, 200 billion+ tokens, over two weeks. The orchestration hardware is disclosed too: two 8-core on-demand CPU nodes ran every orchestrator and agent, each supporting 100+ concurrent subagents with their respective CPython kernels. Because compilation, type-checking, and diffing would saturate any single machine with dozens of agents running at once, they built utilities letting agents outsource heavy work to Prime Sandboxes for parallelization.
Note the composition of this bill: nearly all tokens went through Prime Inference's own GLM-5.3 endpoint, the sandboxes are Prime's own Prime Sandboxes, and the orchestrator is Prime Agent itself. The marginal cash cost of the whole exercise is essentially the internal transfer price of inference tokens. That is the first key lens for understanding this story, and we will return to it.
Why Rust: four reasons, one bonus
The stated reasons are pragmatic, with no language-war flavor:
- Performance: Prime Agent is a long-running daemon with one worker process per session. Native code without a garbage collector accounts for most of the memory and startup gains;
- Concurrency: a single daemon streams model output, runs tool calls, relays messages between agents, and serves every attached client at once. Rust's
SendandSynctraits let the compiler check which data can move between threads and which can be shared; - Compile-time guarantees: exhaustive enums, ownership, and lifetimes rule out whole classes of bugs before the code runs, and Clippy's pedantic lints add hundreds more checks — "which matters when agents write most of the code";
- A chance to redesign the architecture: the codebase was split into nine crates with a one-way dependency graph enforced by Cargo. The TUI's only internal dependency is the shared types crate. The largest source file went from ~15,000 lines in TypeScript to ~2,500, and files over 5,000 lines went from four to zero.
The side benefits are concrete too: each session runs in its own worker process under a small supervisor, so a failure in one session leaves the others running, and sessions persist on disk so clients can reattach after a restart — the "session crash isolation" the post mentions. Transport, process control, and file locking sit behind platform-specific interfaces, so Windows support came down to implementing those traits rather than rewiring the daemon. The client, the daemon, and every worker compile against the same message types in the shared types crate, so a protocol change is checked everywhere it is used. Model and MCP lists became a runtime-fetched catalog, so new models and plugins ship without a Prime Agent release.
The methodology: the root agent writes no code
The most technically substantial part of the post is the orchestration design. A single root agent divided the rewrite into a topological ordering of tasks — but wrote no product code itself. Its job was to monitor every task, merge finished work, maintain an overview of the rewrite, and work with the humans on priorities and decisions. The stated logic is blunt: keeping the root agent away from code frees it to manage and merge.
Each task moved through a four-agent pipeline:
- Planner: writes the overall machine specification, including the feature design, the ground-truth TypeScript behavior, and the verifier (parity check);
- Implementer: writes the Rust code in a dedicated worktree so features develop in parallel;
- Reviewer: adversarially inspects the pull request, using a different model and a separate context from the implementer, looking for reasons the change is wrong;
- Verifier: compiles and runs the feature's parity checks and tests in a fresh Prime Sandbox.
A failed review or verification sends the feature back to the implementer with the findings; the PR merges only once both pass. The post also reveals they are productizing this "factory of finite state machines" pattern inside Prime Agent, so workflows like this can be defined once, reused, and updated.
Two design judgments here deserve underlining. First, the agents that write code and the agents that review it must be separated — in the post's words, "an agent that writes code is biased when evaluating it." That is the human code-review org chart ported verbatim into agent orchestration. Second, the reviewer deliberately uses a different model, not the same model with a different prompt. Same-model correlation bias in review is a real risk; hedging with heterogeneous models is a cheap, effective countermeasure.
Differential testing: letting agents measure their own progress
By the post's own account, the human work in this rewrite was "setting up proper verification to enable autonomous deployment at scale." Parity was split into four specifications, each with an objective check:
- TUI parity: a differential test suite runs the TypeScript and Rust binaries side by side against the same scripted model and diffs the terminal frames each renders. It covers launch, slash menus, tool calls, compaction, the agents view, session resume, subagents, and crash recovery;
- Harness parity: the same run diffs session transcripts and the requests each binary sends to the model provider — same inputs must produce the same sessions;
- Protocol parity: every message type in the daemon protocol is checked against the TypeScript implementation, keeping the ACP and daemon protocol APIs consistent;
- Feature parity: since scripted flows cannot cover an entire interface, agents audited the TypeScript product component by component, classifying each as matching, partial, or missing.
With an objective check for every kind of parity, agents could measure their own progress and catch regressions before merging, which let the team cut back on human review. That sentence is the methodological core of the whole post: the precondition for agent autonomy is not a stronger model, but a verification system that lets agents objectively measure their own progress. Without it, 2,000 agents are 2,000 random-number generators.
The performance numbers: 14× is real, and the caveat is official
Once parity was achieved, a second orchestrator ran a three-day hillclimbing loop with a single objective: keep improving the benchmarks without breaking parity. The procedure is worth describing: agents profile the benchmarks to find where time and memory go, turning the largest costs into a backlog of hypotheses; workers build the current code and the candidate change on the same sandbox and run them in alternating order, alongside neighboring benchmarks to catch regressions elsewhere; two reviewer agents, each on a different frontier model, check that behavior still matches TypeScript, that output stays byte-identical wherever the change claims it, and that nothing regresses on the model-facing surface; the change merges, and the next experiment measures from the new baseline.
One detail shows real craft: they deliberately gave the loop no numeric targets, since a fixed threshold tends to become a stopping point. The agents' only objective was to keep improving the benchmarks without breaking parity for as long as measurable gains remained. Over the cycle, the loop logged 144 experiment and audit records and merged 69 valid changes. Most gains came from three kinds of change: moving work off the startup and render paths, replacing polling loops with event-driven waits, and releasing memory as soon as large sessions finished loading.
The headline numbers from the official benchmark (against their own TypeScript version):
| Metric | Improvement |
|---|---|
| Cold start to input-ready | 14.18× faster |
| Warm start to input-ready | 13.34× faster |
| Large-session memory footprint | 4.76× smaller |
| Installed size | 2.89× smaller |
| Agents view switch latency | 6.09× faster |
The post also includes a comparison table benchmarking the Rust build against external harnesses such as Claude Code, Codex CLI, Pi, and Hermes Agent. Here the official caveat must be quoted verbatim: "Without a common benchmark standard, comparisons should be interpreted with caution." Benchmarks ran on a fresh 4-core, 8 GB Prime Sandbox per run, with noise checks withholding any result too unstable to compare; agents ran in a real terminal driven through a screen emulator against a scripted model, so timings reflect what a user sees and exclude inference. There is no common benchmark standard across products — the caution is the company's own, and it should not be dropped in retelling.
The most honest paragraph: bugs found by dogfooding
Had the post stopped at "14× faster," it would be an ordinary release note. But Prime Intellect wrote the following paragraph, and it is the most valuable one in the piece:
The parity checks only verified the behavior they exercised. Once parity on the main RLM loop was achieved, we moved all our rewrite agents onto Rust, allowing us to dogfood the Rust version at scale — and now Prime Agent Rust was recursively improving itself. Later, we moved our internal team onto the Rust build for daily use, which exposed bugs and missing behavior outside the coverage of the differential tests.
In plain language: passing scripted parity tests does not mean production-ready. However rigorous, automated parity can only verify the behavior it actually exercises; real users — even internal ones — always take paths more adversarial than scripted flows. What followed took weeks of follow-up work: finishing feature ports, fixing bugs, improving performance, polishing the interface. Prime Agent still wrote and tested these changes itself, but humans stayed in the loop throughout — finding problems, directing the changes, and reviewing all results.
One mechanism here is underappreciated: during the dogfood cycles, they had agents review logs and traces from all beta users, automatically diagnosing and resolving runtime issues users had surfaced. The verification system, in other words, is not just "differential tests before writing code" but also "real-traffic replay after shipping." Together they form the complete loop of agent autonomy: measurable before the fact, traceable after it.
For an industry reference point: Microsoft's agent-driven port of the Copilot runtime to Rust (reported by The Register in September) cost roughly $120,000 in tokens plus about three weeks of developer time, converting 430,000 lines of TypeScript into 800,000 lines of Rust — at the price of a few dozen regressions, and with a deliberately conservative module-by-module strategy that left the architecture untouched. Prime Intellect went the other way: two weeks, 2,000 agents, architecture rewritten along the way, and then an admission that "automated checks have blind spots, so we spent weeks of human-in-the-loop polishing." Neither route is strictly better, but writing the blind spots down at least tells the next team how to budget: the cost of agent-written rewrites is not just the token bill — it includes the part parity tests cannot cover, which for now still has to be filled by dogfooding and humans.
The business read: a large-scale ad for its own infrastructure
Strip away the technical narrative and look at the business: what does Prime Intellect sell? Compute, inference (Prime Inference), sandboxes (Prime Sandboxes), eval infrastructure. Every core resource this rewrite consumed — 2,000+ orchestrated agents, 10,000+ sandboxes, 200 billion tokens of GLM-5.3 inference — is its own product line. The blog post is itself a live demo: our sandboxes withstood 2,000 agents compiling and testing concurrently, our inference endpoint fed 200 billion tokens, our agent orchestration completed a full language port.
That is not a criticism. The best marketing an infrastructure company can do is always "we are our own heaviest user." AWS's re:Invent keynotes recount how Amazon retail runs on AWS; Cloudflare describes how it absorbs attacks against itself. What Prime Intellect did here was turn "eating our own dogfood" into a reproducible, auditable engineering report: token counts to two decimal places, agent and sandbox counts, benchmark methodology, caveats — all public. Prospective customers will do their own math: infrastructure that survives this intensity will survive my workload.
One level deeper, this lays track for pricing power in the "agent infrastructure" category. While everyone debates model capabilities, Prime Intellect is shining the light on orchestration, sandboxes, evals, and differential testing — the unglamorous work that is actually the bottleneck when agents move from demo to production. Whoever defines the solution to the bottleneck defines the bill.
An overlooked signal: 200 billion tokens burned on GLM-5.3
One detail deserves to be pulled out on its own: the 200 billion+ tokens went to Prime Inference's GLM-5.3 endpoint — an open-weights model, not the priciest tier of frontier closed models. Combined with the review stage's "different frontier model" arrangement, it reveals Prime Intellect's model staffing strategy: a cheap workhorse model for massively parallel implementation, strong frontier models as a heterogeneous hedge in review.
This staffing is decisive for the cost structure. Had 200 billion tokens been billed at top-tier frontier prices, the rewrite's invoice would be a different order of magnitude. Running implementation on an open-weights model like GLM-5.3 and reserving frontier models for review and key decisions decouples "scale" from "quality": scale comes from cheap tokens, quality is guarded by heterogeneous models at the review gate. It is the same separation-of-verification-and-implementation logic from their org design, projected onto model selection.
For independent developers, this is directly copyable homework: cost-optimizing an agent swarm is not about finding "the cheapest model for everything," but about staffing different tiers of models to different stages. High-concurrency, verifiable stages — writing code, running tests, doing diffs — get the value model; low-concurrency, high-risk stages — architecture decisions, adversarial review, final merges — get the strongest model. Prime Intellect validated the staffing with a 200-billion-token bill. And that number is exactly what they wanted you to see.
What solo teams should steal: the org design of an agent swarm
For indie developers and one-person teams, the most stealable part of this post is not Rust — it is the org design:
First, the root agent writes no code. Many people running agent swarms let one "commander" both split tasks and write code, and its context gets drowned in implementation details while the management function exists in name only. Prime Intellect's approach is counterintuitive but correct: the manager only decomposes, monitors, merges, and decides — never touching code. Context budget is the scarcest resource; a manager's context should be reserved for the global picture.
Second, verification and implementation must be separated — and preferably heterogeneous. An agent reviewing its own code is biased, the same way a human self-reviewing their own PR is no review at all. Going one step further, using a different model for review introduces statistical independence — a genuine "second pair of eyes." For a solo team running agents, the cheapest quality lever available is: implementation and review must never be the same agent, the same model, or the same context.
Third, build the measurement before claiming autonomy. Differential testing — frame-by-frame TUI diffs, transcript diffs, provider-request diffs, protocol message cross-checks — is fundamentally an answer to one question: how does the agent know it got it right? Without objective, repeatable measurement, more agents just mean faster entropy growth. Before scaling up, ask yourself: if the agent ran 100 times, by what standard would I judge this run correct? If you cannot answer, do not scale.
Fourth, budget for the blind spots. Passing scripted parity tests ≠ production-ready — a lesson Prime Intellect paid for with two weeks and 2,000 agents. More practically for indie developers: the real-user paths your tests cannot cover get covered by beta users and dogfooding. Make "have agents audit real logs after shipping" a standing step, not something you remember after an incident.
Ending: the workflow itself became the product
The post closes with Prime Intellect turning the multi-agent workflow behind the rewrite into product capability for users, and weaving Prime Agent deeper into its full stack — cloud agent swarms, inference, traces, sandboxes, evals, hosted training. The Rust release brings Windows beta support and Homebrew installation. The project remains open source.
One-line summary: this was never a "language switch" release; it was a "here is what our infrastructure proved it can survive" release. The 14× figure will age, the benchmark caveat will be forgotten, but three methodological lines will stay: the root agent writes no code, verification is separated from implementation, and the blind spots of automated parity are admitted. For indie developers in the vibe-coding era, the right way to read this post is not "Rust is fast" but "so this is how you organize an agent swarm" — which is, of course, exactly what Prime Intellect wanted you to take away.
Sources
Related articles

Zhipu's flagship GLM 5.3 is generally available on Amazon Bedrock: a 753B-parameter MoE with a million-token context window and a leading 84.5 on the CyberGym security benchmark. AWS demoed it driving the open-source pentest agent Strix in an authorized security test. Behind the listing sits a revenue-share deal on invocation volume — Chinese and American models sold on the same cloud shelf, with Zhipu's Hong Kong shares jumping over 7% on the news.

Between Sept 28 and Oct 4, independent researchers at Swarmchasers found 2,048 public urlquery.net scan reports targeting Alibaba's Amap, peaking at 1,810 in a single day. The agents ran on Tencent Cloud behind a proxy named hysandbox-ats, labeled themselves 'claude' while code fingerprints pointed to Hunyuan and GLM models, and kept no coordination channel at all. This is why the researchers insist on 'fleet,' not 'swarm' — and why agent observability cuts both ways.

AWS has open-sourced Strands Box, a Rust agent sandbox under Apache 2.0 that pairs OS-level isolation with the Dogwood temporal policy engine. Rules can now decide based on the agent's history — cut the network after sensitive files are read, cap Slack posts at three per ten minutes; secrets are injected at the egress gateway so the agent never sees them. macOS first, in developer preview.