Back to Explore
NewsVibeFix 编辑部Updated Oct 11, 2026

OpenAI Finds Its Agents a Sparring Partner: Ironclad's Contracting Workflows Become 11 Training Tasks, Astra Beats Sol by 32%

OpenAI's blog 'Advancing computer use with Ironclad' marks a paradigm shift: computer use moves from general capability to per-application customized RL training. GPT-6 Astra, the first frontier model trained on Ironclad tasks, scored 55.0% vs 41.6% for GPT-5.6 Sol on 11 contracting tasks (8-50 scoring criteria each), with time per attempt falling from 37.0 to 19.2 minutes. The moat is moving from models to the partner list - and OpenAI is openly recruiting the next batch of software companies.

OpenAI and Ironclad partner to train GPT-6 Astra on 11 scored contracting tasks: 55.0% vs 41.6% and 19.2 vs 37.0 minutes against GPT-5.6 Sol

The Sparring-Partner Era: The First "Real Job" OpenAI Gave Its Agents Is Contracts

The takeaway first: OpenAI's computer-use capability has officially shifted from "general capability showcase" to "per-application customized training." On October 6 (see the date note at the end), OpenAI's official blog published Advancing computer use with Ironclad, and the opening line sets the tone: "Our first partner is Ironclad, a leader in AI contracting." The first partner is Ironclad, the contract-management software company. And GPT-6 Astra is "our first frontier model trained on Ironclad tasks" — the first frontier model trained on real Ironclad tasks.

Why does this deserve a deep-dive? Because it marks a turning point in how agents are trained: going forward, it may no longer be "one general model learns to use all software," but "every professional software deserves its own dedicated training." OpenAI said it outright: "We're inviting a small number of software companies to work directly with our research and engineering teams." Note — this isn't an API-integration partnership. It's a training partnership.

11 Training Tasks, 8 to 50 Scoring Criteria Each

What does the training actually look like? OpenAI's recipe is strikingly concrete: "Ironclad employees and people who use Ironclad at OpenAI helped our researchers identify 11 tasks across legal, commercial, and procurement work." Ironclad's own people, plus OpenAI employees who use Ironclad day to day, picked 11 tasks from legal, commercial, and procurement work: setting up an NDA workflow, building a procurement approval flow, updating a reusable legal clause to reflect the jurisdiction a requester selects — the kind of work legal and procurement teams do every day.

The number 11 isn't the point. The scoring density is: "We evaluated each task against 8 to 50 criteria" — depending on complexity. That's the most expensive part of RL training: the reward signal isn't a binary "did it finish?", but a fine-grained rubric written by domain experts. Did it get the jurisdiction right? Are all the approval-chain nodes there? Is the NDA's effective-date logic correct? Eight to fifty criteria, each one a grading event, with the model measured against that ruler at every step.

This is my judgment: the real headline here isn't "OpenAI found a partner" — it's "the raw materials of RL training are on the table for the first time." For the past year, the agent world has talked about RL constantly, but the public conversation was almost entirely about algorithms and compute. What actually determines the capability ceiling — real software environments, expert-designed tasks, fine-grained scoring rubrics — stayed hidden inside labs. This time OpenAI published the recipe: 11 tasks, 8–50 criteria each, experts writing the questions. The next time someone asks "why does my agent keep failing on professional software," check those three ingredients first and see which one is missing.

Astra vs Sol 对比示意图:55.0% 对 41.6% 得分,37.0 分钟降至 19.2 分钟

The Numbers: 55.0% vs. 41.6%, 19.2 Minutes vs. 37.0 Minutes

The scorecard reads: "Astra's average score was 55.0%, compared with 41.6% for GPT‑5.6 Sol, while estimated average time per attempt fell from 37.0 minutes for Sol to 19.2 minutes for Astra." Average score rose from 41.6% to 55.0% — OpenAI's own math calls it 32% higher; estimated time per attempt fell from 37.0 minutes to 19.2 minutes, 48% lower. The comparison is apples-to-apples in the way that matters: Astra at Max reasoning, Sol at High reasoning — each at the setting where it scored best.

Run the plain math: scores up by a third, time cut in half. Better and faster, at the same time — which is unusual in agent benchmarks. Normally "faster" means "sloppier"; score and speed sit on opposite ends of a seesaw. Winning on both ends means the gain didn't come from "thinking fewer steps" but from "genuinely knowing the job." That is the fundamental difference between customized training and a general model: a general model is reasoning; a customized model is practiced. A practiced hand is fast and good — that's just how it works.

One number is easy to skim past: "An internal model used in the development of Astra achieved an even stronger 63.7% on these tasks." An internal model used during Astra's development scored 63.7% on the same 11 tasks. OpenAI didn't say what that model is, but the signal is clear: this rubric still has headroom; 55% isn't the ceiling. My read: 63.7% is written for the next batch of partners — look, there's plenty of road left, plenty of cake left, come aboard.

Honesty footnote first, in OpenAI's own words: every time figure is an "estimated time per attempt" — a simulated estimate, not time real customers actually saved. Secondary sources flagged this explicitly too: the times are simulated from assumed model processing and generation speeds. So when you see 48%, mentally convert it to "relative speedup under lab conditions," not "legal departments can cut headcount in half next year." The numbers are real; the methodology is simulated — every "faster" in this piece carries that qualifier.

Between 55% and 41.6%, the Gap Isn't "Can It Use Software"

55.0% vs. 41.6% — only 13.4 points apart. Sounds small? Flip it around: those 13.4 points buy "understanding the business rules." Sol isn't failing to click Ironclad's buttons — general computer-use capability taught models "look at the screen, click the button, fill the form" long ago. Sol loses its points on things like "wrong jurisdiction selected," "an approval-chain node missing," "stale clause version used." Translation: it can operate; it doesn't understand the business.

That is exactly where enterprise agents have hurt most over the past year. In the demo, the agent flies through the CRM and ERP; in production, it crashes — not because it clicked the wrong button, but because it violated a business rule: it sent a discount it shouldn't have, skipped a critical approval node, left a compliance hole in a contract. Between "can operate the software" and "can do the job" lies an entire set of tacit business rules, and those rules live in veterans' heads and company wikis — never in training data. What Ironclad's 11 tasks do is make the tacit explicit: 8 to 50 scoring criteria are, in essence, an operations manual for "what doing this job correctly actually means."

So my judgment: when evaluating an enterprise agent, don't start with "how many apps does it support" — start with "how many scored tasks has it done per app." The first is a marketing number; the second is a capability number. OpenAI just published a new capability currency: task count × rubric density. Ironclad is 11 × (8–50). What will the next partner's number be? That product is about to become the arms-race metric every agent vendor tracks quietly.

RL 训练闭环示意图:真实软件环境与 8~50 条细粒度评分标准

The Moat Moved: From Models to the "Partner List"

Zoom out. The most thought-provoking line in the post isn't a number — it's this: "We're inviting a small number of software companies to work directly with our research and engineering teams." Translation: OpenAI is openly recruiting the next batch of sparring partners. Limited seats ("small number"), and the collaboration depth is "work directly with our research and engineering teams."

Think about what that implies. The agent moat used to sit in the model: whoever had the smarter model won. Now OpenAI has moved the moat down a layer, to "who holds real software environments + expert tasks + scoring rubrics." Models can be caught up with; compute can be bought. But "Ironclad's legal experts sat down and wrote you 11 tasks and hundreds of scoring criteria" — that kind of partnership resource is exclusive, slow, and irreproducible. Today it's Ironclad; tomorrow it could be Salesforce, ServiceNow, Workday… Every name added to the list "customizes away" agent capability for one more category of professional work.

This is a direct signal for vertical-SaaS readers: the window is open now, and it's the "small number" limited edition. If your software has a complex professional workflow (legal, finance, healthcare, engineering) and you don't have an agent strategy yet — OpenAI is looking for you. And flip it around: if your competitor gets on that list first, the agent version of "the power user of X software" grows up in their house first. Once this kind of partnership starts, the data flywheel (tasks → scores → training → better agents → more user behavior) tilts toward the partner. First aboard and late aboard are not eating the same dividend.

One more connection: our earlier piece on OpenAI's Decisions API was about judgment models solving "what to choose"; this Ironclad partnership solves "how to do it." Judgment layer + execution layer — OpenAI is placing pieces on both layers of agent infrastructure at once: the Decisions API turns "making decisions" into a standard part; the Ironclad partnership turns "doing professional work" into something trainable, scorable, and replicable. Read together, OpenAI's agent roadmap is now legible: general base + per-application customization + per-step standard parts.

The Signal for Developers: Stop Competing on Prompts Alone

Something you can act on tonight. If you're building an agent for a professional software product — your own SaaS or a client's internal system — this post hands you homework you can copy directly:

  • Find your "question writers" before tuning the model. Step one of the Ironclad playbook isn't "switch to a bigger model" — it's "find the people who know this job best and have them write 11 tasks." The most senior domain expert on your team is worth more than any prompt engineer. Have them write down, line by line, "what counts as doing this job right" — that's your rubric, your reward signal.
  • Score down to the business-rule level, not the task-complete level. The 8-to-50 density is your reference line: "did the form submit?" is one criterion; "jurisdiction, approval chain, clause version, effective-date logic" is thirty. Agents fail on professional software at the second level, always.
  • Put time in the eval too. OpenAI reported both the score and "estimated time per attempt" — and reported it even at simulated fidelity. In an enterprise agent's ROI story, "faster and better" are two independent dimensions — accuracy alone leaves business stakeholders unmoved.
  • Decide whether you want to be a sparring partner. If you're a vertical-SaaS vendor, OpenAI's invitation is a two-way deal: you bring the environment, the experts, and the tasks; in return, your software becomes natively fluent on the strongest models. Worth doing the math — the window won't stay open forever.

The Skepticism, Stated Upfront

Applause aside, four hard objections.

First, the people who wrote the questions and the questions used for grading are the same set. FourWeekMBA's critique cuts to the bone: Astra is "trained on Ironclad tasks," and it's tested on "11 research tasks chosen with the same partner" — training and testing share a source. Of that 32% gain, how much is "genuinely learned contracting work" and how much is "memorized these 11 tasks"? OpenAI ran no cross-software, cross-contract-type generalization tests. Until we see "the model trained on Ironclad also performs on someone else's contracting software," discount the number.

Second, the time figures are simulated — covered above, no need to repeat, but it must be nailed down once more: 48% lower is "estimated," not measured from customers. Anyone pasting that number straight into an ROI deck is passing off lab data as production data.

Third, 55% is still low in absolute terms. Eleven tasks, experts hand-writing the questions, hundreds of criteria fed in — and a frontier model still lands at 55%. Contracts are high-stakes: one wrong jurisdiction can mean a lawsuit. 55% means "wrong roughly every other task" — nowhere near "unsupervised," though close to "useful under supervision." Read the absolute value, not just the relative gain.

Fourth, "small number" is itself the problem. OpenAI's list is finite and first-come. That turns agent "professional capability" into an allocated resource: whoever partners first gets their software "custom-trained" first. Small vendors, open-source tools, long-tail industry software may never get on the train. As the moat moves from models to the partner list, the allocation of capability moves from a technical question to a relationship question — which isn't necessarily good for the ecosystem.

But none of the four shakes my core judgment: the paradigm is real; the numbers are discounted. The second half of computer use isn't "a more general model" — it's "more specialized training." Building agents for professional software is no longer a prompt-engineering job; it's a three-piece job: real environment + expert tasks + fine-grained scoring. Ironclad is the first domino. The next name on the list is the one truly worth watching.

Primary source: OpenAI's official blog, Advancing computer use with Ironclad. Date note, stated transparently: the blog post body carries no visible date stamp; four independent secondary sources (mediarelease.co, fourweekmba, bytewatchr, subagentic) plus OpenAI's official LinkedIn post consistently date it October 6, 2026, which is the date used here. Cross-checked against FourWeekMBA and Subagentic.

Sources

Browse projectsPublish your project

Related articles

Diagram of Claude dynamic workflows: a lead agent writes a program that fans out to hundreds of subagents in phases and merges their results
News
Anthropic Ships Agent Orchestration: Dynamic Workflows in Public Beta, 1,000 Subagents per Run

Dynamic workflows for Claude Managed Agents entered public beta on October 9: an agent writes a program that runs many agents in phases and merges results — max 1,000 agents per run, 64 at once, 24-hour default lifetime. The launch rebuts the 'waste of tokens' charge with a head-to-head: 70 planted bugs, a single agent found 14/15/27 across three runs, a workflow found 66 every time. Billing: tokens at each model's rates plus $0.08/session-hour. Token math and a use-it-or-save-it checklist.

Product LaunchAI CodingDeveloper Workflow
Illustration of the Nemotron dual-gold recipe: SFT and RL checkpoints, 22000 curated programming problems, and the GenCorrect generate-evaluate-refine inference loop
News
NVIDIA Open-Sources Its 'Double Gold' Training Recipe: Nemotron Beats Top Human IOI Score, Full 22,000-Problem Dataset Released

NVIDIA's Nemotron systems hit gold level at IOI 2026 (535.4/600, above the top human score of 498.27 in an unofficial run) and IMO 2026 (30/42, graded by official IMO graders) — and the team open-sourced the full recipe: SFT/RL checkpoints, both training datasets, a new 200-problem olympiad benchmark, inference pipelines, and prompts. The lesson is co-design of model, data, and inference loop: GenCorrect's generate-evaluate-refine cycle carried a 291-point model past the 438.3 gold bar.

Open-source ProjectsModel UpdatesAI Coding
Concept art for Manus 2.0 and personal agent Cue: work-creation and personal-life tracks
News
Manus Is Back: Butterfly Effect Raises Over $500M at ~$4B Valuation

Butterfly Effect raised over $500M led by Boyu and IDG, with Tencent, Sequoia China, and ZhenFund joining, targeting a ~$4B valuation. One month after going independent following the collapsed $2B Meta deal, Manus 2.0 and personal agent Cue launched — a comeback for China's general agent.

Product NewsStartup JourneyIndustry Trends