Anthropic Ships Agent Orchestration: Dynamic Workflows in Public Beta, 1,000 Subagents per Run
Dynamic workflows for Claude Managed Agents entered public beta on October 9: an agent writes a program that runs many agents in phases and merges results — max 1,000 agents per run, 64 at once, 24-hour default lifetime. The launch rebuts the 'waste of tokens' charge with a head-to-head: 70 planted bugs, a single agent found 14/15/27 across three runs, a workflow found 66 every time. Billing: tokens at each model's rates plus $0.08/session-hour. Token math and a use-it-or-save-it checklist.

A Launch With a Benchmark Attached: Anthropic Absorbs Agent Orchestration
On October 9, Anthropic moved dynamic workflows for Claude Managed Agents into public beta. The release notes define it in one line: "A workflow is a program that runs many agents in phases and combines their results." A workflow is a program that runs many agents in phases and combines their results.
But the function isn't the most interesting part of this story — the gunpowder is. Not long ago, a senior OpenAI engineer publicly called agent swarms a waste of tokens. Anthropic's response was characteristically Anthropic: no flame war, just a head-to-head benchmark. They planted 70 bugs in a 116k-line codebase. A single agent found 14, 15, and 27 across three runs; a workflow found 66 in each of its three runs. 14–27 vs. 66 — that is Anthropic's answer to the "waste of tokens" charge.
The verdict up front: this is a key step in multi-agent orchestration going from hand-rolled craft to official platform capability. For the past year, fan-out logic, concurrency management, result merging, and retries have been scaffolding every team hand-built — in LangChain, in CrewAI, in countless indie developers' late-night code. Now Anthropic has absorbed all of it into the platform: you describe the task, the agent writes the orchestration program, and the server runs it in the background. For indie developers, the decision point isn't "is it cool" — it's the token economics: 64 concurrent agents, a 24-hour run lifetime, $0.08 per session-hour. When to use it, when to save it. This piece does the math.
What a Dynamic Workflow Actually Is: A Program Written by the Agent, Not by You
Nail down the key distinction first. Managed Agents already had subagents: the lead agent delegates to child agents and can follow up with each one. Dynamic workflows are the second mode — the lead agent writes a program (the workflow) that decides for itself how many phases to run, how many agents to start, how results pass between them, and what to do on failure, and then the server executes it in the background. The main thread stays free to keep talking to you.
There is no "start a workflow" API call. You describe the work in your message or the system prompt, and the agent decides whether and when to start a run. The run moves through named phases; each agent works in its own session thread, results flow back into the program, and the program passes them to the next agent. The trade is control: once a run starts, the agent cannot send follow-up messages to its threads — the program runs its written logic to the end, and the server archives each thread when the run finishes.
This "agent writes the program, server runs the program" structure is the design choice worth savoring. Old multi-agent frameworks were "you write the orchestration code, the model fills in the content." Now it's inverted — the orchestration logic itself is model-generated. How to slice the phases, how to divide labor among agents, how to merge results — architecture decisions a human used to make are now made by the agent at runtime. That's what "dynamic" means: not executing a DAG you hard-coded, but letting the agent write the DAG on the spot.
The Limits: 1,000, 64, 24 Hours, 10
The docs state the caps plainly, quoted verbatim: "A run can start up to 1,000 agents over its life, with up to 64 working at once, a 24-hour default lifetime and up to 10 runs open per session by default." In other words: one run can start at most 1,000 agents over its lifetime, with at most 64 working at once, a 24-hour default lifetime, and at most 10 runs open per session by default.
A few details worth noting. First, 1,000 is the count of agents started, not threads — a failed agent gets rerun on a new thread, so thread count can exceed 1,000, and going over the cap ends the run with a thread limit error. Second, 64 concurrent is a ceiling, not a guarantee; the docs say so explicitly. Third, runs don't nest: only the agent on the session's primary thread can start a run, so an agent inside a run can't start one of its own. Fourth, 24 hours is the default lifetime and the agent can set it shorter — but a pause (for example at the budget cap) doesn't pause the lifetime clock.
Switching it on is a one-line config change: "Set the agent's multiagent type to multiagent_20261001 with the managed-agents-2026-04-01 beta header." Set the agent's multiagent type to multiagent_20261001 with the managed-agents-2026-04-01 beta header. With that type, workflows are on by default — then you tell the agent in the system prompt when to start a run, and follow progress through workflow_run.* events on the session's event stream.
My judgment: these parameters are designed for overnight batch jobs, not real-time interaction. A 24-hour lifetime, background execution, event-stream tracking — Anthropic is imagining you dropping a "find the change-of-control clauses in these 300 contracts" task before leaving work and collecting results in the morning. The docs' own example is exactly that 300-contract review. Which explains the control trade-off: long, wide-parallel tasks were never meant to be watched in real time.
The 70-Bug Benchmark: The Real Signal Is in the Variance
The official claim: "We planted 70 bugs in a 116k-line codebase. Across 3 runs, a single agent found 14, 15 and 27 bugs. A workflow consistently found 66 in each of its 3 runs."
Do the percentages first: the single agent hit 20%, 21%, and 39% across its three runs; the workflow hit 94% all three times. The mean gap is more than 3x (roughly 19 vs. 66), but the real signal is in the variance: the single agent's best run (27) nearly doubled its worst (14), while the workflow returned the identical 66 three times. For production, that "three identical runs" is worth more than the number 66 itself — repeatability is the admission ticket to production. Nobody puts a system that scores 94% once and 40% the next into CI.
Why is the workflow so much steadier? The mechanism isn't mysterious: a single agent is "one person reading cover to cover," with attention, luck, and context window all fighting 116k lines of code; a workflow is "slice the codebase into dozens or hundreds of chunks, assign each chunk to a dedicated agent for a careful read, then merge the findings." It's the classic playbook of trading parallelism for coverage — in code audit, the human-wave tactic wins not on IQ but on engineering discipline: every line actually gets read carefully.
But three question marks hang over this benchmark, and Anthropic hasn't hidden them. First, it wrote the test and graded it: its own codebase, its own planted bugs, no independent third-party replication. Second, no token bill was attached: how many tokens did those 66 bugs burn, and at what dollar cost? The official text says nothing — and that's precisely the core of the "waste of tokens" debate. Third, finding bugs isn't fixing bugs; the test measured detection, not repair. So my position: the directional evidence holds; don't treat the numbers as scripture. Until an independent team reproduces this on its own codebase, put an asterisk in front of 94%.
The Cost Math: The $1.92 Isn't the Number That Matters
The official billing line: "its agents' tokens bill like the rest of the session, at each model's rates, and Managed Agents adds session runtime at $0.08 per session-hour." A run has no price of its own: its agents' tokens bill like the rest of the session at each model's rates, and Managed Agents adds session runtime at $0.08 per session-hour, counted only while the session is running.
Start with the fixed math: a full 24-hour run costs 24 × 0.08 = $1.92 in session runtime. Under two dollars — the infrastructure slice is negligible. The real bill is tokens: 1,000 agents, each with its own system prompt, task context, and several rounds of tool calls — token usage scales linearly with agent count. Anthropic admits as much: "Dynamic workflows are powerful and can use a lot of tokens, so we suggest starting with a scoped task." Powerful, and token-hungry — start with a scoped task.
Here's the easy misread: $0.08/session-hour looks like "longer runs cost more," but for a 64-concurrent run, time is actually the money-saving dimension — parallelism compresses wall-clock time, which makes session runtime cheaper; what's expensive is how many agents you dispatched and how much context each one read. So the cost lever isn't "run faster," it's "dispatch fewer agents with smaller contexts": the finer you slice the phases and the tighter each agent's scope, the more controllable the token bill. That's the technical meaning of "start with a scoped task" — first measure "tokens per bug" on a small task, then decide whether to scale to the whole repo.
There's also a brake pedal: the session budget. A session can set a spend cap; when it's hit, every open run pauses, and each working agent finishes only the request it already started. The budget is the fuse built for overnight batch jobs — while you sleep, a run can't burn through your credit card. My advice: the first time you run any workflow, set a budget you'd be fine losing, and treat it as tuition for measuring token cost.
My judgment: the cost model is "fixed part negligible, variable part scales linearly." That makes it a natural fit for scenarios where task value ≥ token cost × a safety margin: one critical vulnerability found in a full-repo security audit is worth far more than tens of dollars in tokens. But using it for a "take a quick look at this code" chat you could finish in five minutes is taking 64-way concurrency to swat a fly.
Use It / Don't: A Decision Checklist
Four shapes that fit:
- Repo-scale code audit and security analysis — the official benchmark's scenario, and the best mechanical fit: the task splits cleanly, results merge cleanly, and coverage matters more than single-point brilliance.
- Large migrations and bulk refactors — mechanical changes across hundreds of files, one agent per file, reconciled at the end. Caveat: validate the fix rate on a small scope first, not just the find rate.
- Batch document review — first-pass triage across hundreds of contracts, résumés, or papers. The docs' own example is finding change-of-control clauses in 300 contracts — an officially blessed shape.
- Long tasks with repeatable orchestration — the same pipeline every week: an agent-written workflow is itself a reusable program, more adaptable to input drift than a hand-written script.
Four shapes that don't:
- Real-time interactive tasks — a run is asynchronous and backgrounded; once started you can't send follow-ups, and changing requirements mid-run means stopping and restarting. Don't use it where you need to iterate conversationally.
- Small tasks — anything a single agent finishes in minutes doesn't justify the orchestration overhead (writing the program, phasing, merging). "Start with a scoped task" is a ceiling recommendation, not a floor.
- Tasks needing fine-grained control of every step — after a run starts, control passes to the program; you can watch the event stream but can't intervene. For steps with hard compliance requirements, wait.
- Budget-sensitive exploration — a "let's just run it and see" mindset paired with a 1,000-agent ceiling is the fastest way to burn money. Explore with a single agent first; promote to a workflow once the shape is validated.
The one-line version: the bigger, the more splittable, and the more wait-tolerant the task, the more a workflow pays off; the smaller, the more interactive, and the more control-sensitive, the more a single agent wins. A dedup note while we're here: we previously covered GitHub Copilot's Dynamic Workflows — that's a different capability inside GitHub's product. This is Anthropic's infrastructure bet at its own platform API layer. Same lane, different vendor.
The Platformization Signal: Orchestration Goes From Hand-Rolled to Utility
Zoom out. The real historical position of this launch: multi-agent orchestration is sinking from framework-layer craft to platform-layer utility. For two years, fan-out, concurrency control, result merging, and failure retries were wheels every agent framework — LangChain, CrewAI, and countless indie developers' late-night code — had to reinvent. Now Anthropic says: stop building them; the platform runs them. You just describe the task and write the scheduling policy in the system prompt.
Once that sinking completes, the competitive landscape above shifts. Framework-layer value moves from "how to get agents running" to "how to describe the task well" — those few lines in the system prompt about when to start a run, how to slice phases, and what to do on failure become the new moat. And the felt change for indie developers is even more direct: you no longer write code for parallelism; you do accounting for parallelism. 64-way concurrency isn't a technical flex — it's a cost unit. 1,000 agents isn't a numbers game — it's a budget ceiling.
My judgment: within 12 months, "workflow-shaped problem" will become the selection jargon of agent teams — first question on any task: is this single-agent-shaped or workflow-shaped? Whoever can answer that reliably will save real money over whoever just cranks up concurrency. And Anthropic's real gambit here is the 66-vs-14–27 benchmark: dragging the "should we go multi-agent" debate from philosophy down to engineering — stop arguing about wasted tokens; run it once and let the numbers talk.
The Skepticism, Stated Upfront
Applause aside, three hard objections.
First, the benchmark is home-grown. Seventy bugs planted in its own 116k-line codebase — the planter and the finder are the same company. The bug distribution, the difficulty curve, the code style: all shapes Anthropic knows intimately. Until independent teams reproduce it on their own codebases, discount the 66 — it proves the mechanism works, not that 94% is universal.
Second, the most important number is missing: the token bill. The whole debate started over wasted tokens; Anthropic answered the "effectiveness" half and skipped the "cost" half. How many tokens did 66 bugs cost? What's the per-bug token price? Without those numbers, the "use it or save it" math never closes. The second half of this debate begins when someone publishes the first independent "tokens per bug" report.
Third, beta terms can change anytime. The 64-concurrency figure is explicitly not guaranteed; the 1,000 / 24-hour / 10-run defaults and the $0.08 pricing are all public-beta numbers. Tighter limits or repricing at GA wouldn't surprise anyone — don't bake beta parameters into your long-term cost model. Also note: this is a platform API capability, not a toggle in the Claude chat app. Don't go looking in the app settings.
But none of the three shakes the core judgment: the platformization of orchestration is real; the numbers are pending verification. For indie developers, what matters isn't whether 66 is precise to the decimal — it's that as of today, "dispatching 1,000 agents at once" went from "three months of hand-rolling" to "one line of config." Run your own cost math, train your own shape judgment — but Anthropic just moved your starting line forward.
The agent race is shifting from "how smart is one model" to "how much smart can a platform schedule." And this time, Anthropic proved something with 70 bugs: on some tasks, the discipline of 1,000 ordinary agents beats the inspiration of one brilliant one.
Primary source: Anthropic's official Workflow runs documentation (October 9, 2026 dynamic workflows public beta release-notes entry, @ClaudeDevs launch thread at 16:12 UTC, multi-agent orchestration and pricing pages in the same docs; third-party roundup CellCog quotes the official texts above).
Sources
Related articles

OpenAI's blog 'Advancing computer use with Ironclad' marks a paradigm shift: computer use moves from general capability to per-application customized RL training. GPT-6 Astra, the first frontier model trained on Ironclad tasks, scored 55.0% vs 41.6% for GPT-5.6 Sol on 11 contracting tasks (8-50 scoring criteria each), with time per attempt falling from 37.0 to 19.2 minutes. The moat is moving from models to the partner list - and OpenAI is openly recruiting the next batch of software companies.

NVIDIA's Nemotron systems hit gold level at IOI 2026 (535.4/600, above the top human score of 498.27 in an unofficial run) and IMO 2026 (30/42, graded by official IMO graders) — and the team open-sourced the full recipe: SFT/RL checkpoints, both training datasets, a new 200-problem olympiad benchmark, inference pipelines, and prompts. The lesson is co-design of model, data, and inference loop: GenCorrect's generate-evaluate-refine cycle carried a 291-point model past the 438.3 gold bar.

Perplexity shipped Decider, an open-weights decision model: 27B params, Apache 2.0, V1 at $0.04/M input — cut to $0.02/M six days later in V1.1. Self-reported 85.71% vs Jev's 84.51%: the closed-vs-open script replays in a new lane.