Back to Explore
NewsVibeFix 编辑部Updated Oct 11, 2026

NVIDIA Open-Sources Its 'Double Gold' Training Recipe: Nemotron Beats Top Human IOI Score, Full 22,000-Problem Dataset Released

NVIDIA's Nemotron systems hit gold level at IOI 2026 (535.4/600, above the top human score of 498.27 in an unofficial run) and IMO 2026 (30/42, graded by official IMO graders) — and the team open-sourced the full recipe: SFT/RL checkpoints, both training datasets, a new 200-problem olympiad benchmark, inference pipelines, and prompts. The lesson is co-design of model, data, and inference loop: GenCorrect's generate-evaluate-refine cycle carried a 291-point model past the 438.3 gold bar.

Illustration of the Nemotron dual-gold recipe: SFT and RL checkpoints, 22000 curated programming problems, and the GenCorrect generate-evaluate-refine inference loop

The Medals Are the Headline; the Recipe Is the Story

Start with the most startling numbers: at IOI 2026, Nemotron-3-Ultra-CC scored 535.4 out of 600. The gold threshold was 361.12, and the top human contestant that year scored 498.27. The machine didn't just clear the gold bar — it left the best human score behind by nearly 37 points. On the other front that same season, the Nemotron system scored 30 out of 42 at IMO 2026, clearing the official gold threshold of 29.

But this is not another "AI beats humans" victory lap. Quite the opposite — if you only stare at the medals, you'll miss what actually matters in this Hugging Face blog post from the NVIDIA team: they open-sourced the entire training recipe. The SFT and RL checkpoints, both training datasets, a brand-new 200-problem olympiad-level math benchmark, the inference pipelines, the prompts, even the proofs submitted at the IMO — all public.

For VibeFix readers, that's the real story. We don't lack "some model broke another record" news. We lack a reusable, complete manual for turning a general foundation model into a domain expert. That's exactly what NVIDIA just handed over.

One transparency note first: the blog page itself carries no explicit date stamp. The "October 7, 2026" date comes from cross-confirmation by multiple outlets including smartchunks and tpsreport — each writing that "NVIDIA published a report on Hugging Face on October 7, 2026." I'm using that date, and telling you plainly where it comes from.

IOI 得分对比示意图:Nemotron 535.4 分超越人类最高分 498.27

Two "Golds," Each Worth Something Different

Start with reading comprehension, because these two golds are not equally "official" — conflating them is clickbait.

The IOI one is a "quasi-gold." The post is admirably honest: "It was an unofficial, unsupervised benchmark and was not included in the official IOI ranking." Translation: this was a live, prospective run where the machine competed under the same time limits, no-internet rule, and submission constraints as the human contestants — but the score was never entered into the official IOI ranking, and there was no on-site supervision. So "beating the top human score of 498.27" is a real numerical comparison, but it is not an official record. NVIDIA put that sentence in the first paragraph under the results table — that honesty deserves credit, and it reminds us: when reading AI competition news, always read the footnotes first.

The IMO one is a real gold. "The IMO system's submitted proofs were graded by official IMO graders." The submitted proofs were scored by official IMO graders: 30 out of 42, above the official 29-point gold threshold, with full credit on four of the six problems. And note that the system "worked in natural language, with no formal prover, external tools, or internet access" — proofs written in plain prose, judged by human grading standards, and awarded gold.

One unofficial benchmark, one officially graded result — this distinction isn't PR spin; it's the key to understanding the whole thing. It tells you: programming competitions are won with a "generate-run-fix" engineering loop, while math olympiads are won with a "generate-critique-refine" reasoning loop. Two loops, one idea — and NVIDIA turned them into a single methodology.

The Four-Step Recipe: Medals Aren't Forged, They're Assembled

The most copy-worthy passage in the post is NVIDIA's own summary of "a reusable specialization recipe." Four steps:

Step one: start from a strong base. Both projects built on the Nemotron 3 family. The programming side used Nemotron-3-Ultra-CC (550 billion total parameters, 55 billion active) and Nemotron-3-Nano-CC (30 billion total, 3 billion active); the math side used Nemotron 3 Ultra. Nobody trained a new foundation model from scratch for each competition — "We did not need to build a new foundation model for every challenge. We specialized Nemotron for the task."

Step two: curate domain problems plus high-quality reasoning traces. The programming side curated 22,000 problems and generated synthetic reasoning traces; the math SFT corpus held 414,890 quality-filtered examples across 15,818 unique proof problems. Note the scale here: not "millions of generic examples," but "fifteen thousand high-quality proof problems, multiple traces each." Density of quality beats volume of data.

Step three: standard post-training — SFT, plus RL where useful. No black magic, just SFT + RL. But there's a counterintuitive finding: for the stronger Ultra model, a single SFT epoch outperformed the fully post-trained (SFT+RL) Nano model — across IOI, ICPC, and LiveCodeBench Pro. The stronger the base, the cheaper the specialization. That's good news for teams on a budget: picking the right base matters more than stacking post-training rounds.

Step four: pair the specialist with a generate-evaluate-improve inference loop. GenCorrect on the programming side, generate-verify-refine on the math side. The IOI 2025 numbers make the case best: the Nano model went from 130 points before post-training to 280 after SFT and 291 after RL — then with GenCorrect switched on, 468 points, straight past the 438.3 gold threshold. Ultra-CC hit 502 with the same test-time strategy. In other words, the inference loop's gain (291→468) was bigger than SFT+RL combined (130→291).

四步炼丹配方流程示意图:基座、数据、SFT/RL、GenCorrect 循环

The Methodology Behind the Soundbite: Co-Design

The post's thesis statement, in my view: "The medals were not produced by fine-tuning alone, and they were not produced by brute-force sampling alone. They came from co-designing the model, the data, and the inference loop."

That sentence calls out two popular lazy doctrines by name. The first is "fine-tuning fixes everything": grab an open model, feed it some domain data, SFT it, and expect an expert. The second is "sampling fixes everything": if the model is weak, just sample more — best-of-N over thousands of draws, and one will pass. NVIDIA's data says neither alone reaches gold. Nano's SFT+RL topped out at 291, still far from the 438.3 gold bar; it was the GenCorrect iteration loop that carried 291 to 468.

The subtler insight comes from the IMO side: combining complementary SFT and RL checkpoints beat brute-force sampling from a single checkpoint. The SFT checkpoint was strongest in the first search round, the RL checkpoint best overall as a single checkpoint — complementary strengths, so the final system deployed the general model plus both specialists together. In engineering terms: diversity beats volume. One that critiques, one that constructs, one as the foundation — that's already a miniature committee.

My judgment: model competition in 2026 is shifting from "train a bigger base" to "design a better inference loop." Base models are commoditizing (even NVIDIA says there's no need to rebuild one per challenge); the real differentiation sits in the data recipe and the test-time system design. This post's value is that it lays the "system design" layer completely bare — pipelines, prompts, and submitted proofs all in the NeMo-Skills repository.

What a One-Person Company Can Steal: GenCorrect Is the Academic Version of a Test-Driven Agent Loop

Now the actionable part. GenCorrect at its core is: generate candidates → evaluate them somehow (run tests, score, critique) → improve against the feedback → repeat. For a one-person company building AI coding workflows, isn't that exactly what you write every day?

Your agent writes code, runs the tests, the tests fail, you feed the error back, it fixes, tests run again — that's the grassroots version of GenCorrect. NVIDIA's version is just more deliberate in three places: first, the evaluator is itself trained (the IMO data covered proof generation, refinement, verification, and meta-verification — the model learned to judge whether a proof was complete); second, the candidates are diverse (multiple complementary checkpoints generating in parallel); third, the refinement is structured (score first, critique next, then refine only the most promising attempts — not blind retries).

So the borrowable checklist is concrete:

  • Treat the "verifier" as a first-class citizen. In most people's agent loops, verification is just "run the tests and see." NVIDIA's approach trains models specifically for verification and critique. The low-cost version for a solo shop: build a solid evaluation set for your core task, and dedicate a separate model instance purely to code review (reviewer and generator with different prompts, or even different models).
  • Diversity over sample count. Instead of drawing 20 samples from the same model, draw a few each from 2–3 differently configured or differently prompted instances, then select. The IMO evidence backs this up.
  • Base-model choice matters more than post-training spend. Ultra with one SFT epoch beat Nano with full post-training — mapped onto your workflow: spend time picking the right flagship model first, not carving prompts. Prompt engineering has a ceiling; base choice sets the floor.
  • Record your reasoning traces. NVIDIA's 22,000 problems and 414,890 proof examples are, at bottom, good problem-solving processes distilled into data. The gnarly bugs your agent successfully fixed in daily work — the errors, the fix attempts, the final solutions — are your domain traces. Accumulate them; they're the dataset for fine-tuning your own coding assistant later.

This isn't motivational fluff. The NeMo-Skills repository contains the full implementation of the IMO inference pipeline and its prompts — you can read exactly how its generate-verify-refine is orchestrated. This is the first time a major lab has put a source-level implementation of a gold-level agent loop on the table. That kind of thing used to exist only as pseudocode in papers.

Cold Water: Four Buckets, Poured One by One

Bucket one: the post itself admits the training and inference runs were "substantial." SFT on a 550-billion-parameter model, plus multi-round GenCorrect inference — that's not reproducible on your laptop. The recipe is public; the compute bill is yours. The right way to read this post is for the design ideas, not the parameter counts.

Bucket two: the IOI score's "unofficial" status caps what it means. An unsupervised run with no on-site supervision is inherently a tier below in competitive seriousness. NVIDIA wrote that honestly; media retellings often swallow the footnote. When you cite "535.4 beating the top human score," always carry the qualifier — "unofficial benchmark, not in the official ranking." That's the dividing line of professionalism.

Bucket three: the IMO 30 points belong to the "system," not the "model." The generate-verify-refine system plus a high-compute selection stage plus three cooperating checkpoints got to 30/42. What does the bare model score on a single generation? The post doesn't say. Don't misread a system score as model capability — that's the flip side of the co-design idea: the capability lives in the system, not in any single checkpoint.

Bucket four: open-sourced doesn't mean ready to run. Checkpoints, datasets, the benchmark, and the pipelines are all on Hugging Face and GitHub — but standing them up and adapting them to your domain is still engineering work. NVIDIA open-sourced the recipe and the ingredients, not a ready meal. That said: before this, the big labs wouldn't even show you the recipe. Going from zero to one here is precious enough.

Why This Deserves a Headline

For two years, AI competition news followed a fixed template: a big lab's model scores well at some contest, the paper goes on arXiv, code is "coming soon." NVIDIA just tore up the template: scores, checkpoints, training data, the benchmark, the inference pipelines, the prompts, the submitted proofs — all public. This isn't PR; it's an infrastructure drop. Those 200 olympiad-level problems in Nemotron-IMO-Bench are, from today, the public examination hall for every math-reasoning model; the synthetic-trace method over 22,000 programming problems is the reference implementation for every coding specialist.

One level deeper, this is NVIDIA playing a long game: positioning Nemotron as "the most specializable base." "Easy to fine-tune should mean more than making a checkpoint trainable" — good fine-tunability shouldn't just mean the checkpoint is trainable, but that a clear, reusable recipe exists. When every team trains its own expert models with Nemotron's recipe, the Nemotron ecosystem becomes the de facto standard. Open-sourcing the recipe is the smartest ecosystem play there is.

For VibeFix readers, my advice is one sentence: go read the IMO inference pipeline in the NeMo-Skills repository. Not to reproduce an IMO gold — but to see what a production-grade generate-evaluate-improve loop looks like in code: how the prompts are written, how critique and refinement divide the labor, how multiple checkpoints cooperate. The missing piece in your agent workflow is very likely hiding in those few hundred lines.

Medals tarnish; recipes don't. What NVIDIA handed over this time isn't a trophy — it's an open cookbook.

Primary source: Hugging Face official blog (NVIDIA team), One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO. The blog page carries no explicit date stamp; the October 7, 2026 publication date is cross-confirmed by multiple outlets including smartchunks and tpsreport ("NVIDIA published a report on Hugging Face on October 7, 2026"). Related papers and code: the IOI / IMO arXiv papers, the NeMo-Skills GitHub repository (IMO inference pipeline, prompts, submitted proofs), and the Nemotron Labs IMO 2026 collection on Hugging Face (SFT/RL checkpoints, both training datasets, Nemotron-IMO-Bench).

Sources

Browse projectsPublish your project

Related articles

Conceptual illustration of Zhipu GLM 5.3 joining the AWS Bedrock model shelf
News
After OpenAI's Agents, AWS Puts Zhipu's GLM-5.3 on the Bedrock Shelf: Chinese and American Models on the Same Cloud

Zhipu's flagship GLM 5.3 is generally available on Amazon Bedrock: a 753B-parameter MoE with a million-token context window and a leading 84.5 on the CyberGym security benchmark. AWS demoed it driving the open-source pentest agent Strix in an authorized security test. Behind the listing sits a revenue-share deal on invocation volume — Chinese and American models sold on the same cloud shelf, with Zhipu's Hong Kong shares jumping over 7% on the news.

Model UpdatesAI CodingIndustry Trends
Illustration of 2,000 AI agents collaboratively rewriting the Prime Agent codebase from TypeScript to Rust
News
2,000 Agents Rewrote Themselves in Rust: Prime Intellect's Two-Week Dogfooding Experiment

Prime Intellect had Prime Agent orchestrate 2,000+ agents to rewrite itself from TypeScript into Rust in two weeks, burning 200B+ tokens across 10,000+ sandboxes. The real story is not the 14x speedup but the honest methodology: a root agent that writes no code, verification separated from implementation, and an open admission that passing scripted parity tests does not mean production-ready.

AI CodingOpen-source ProjectsIndustry Trends