Self-Improvement Itself Needs Regularization: Google's RRSI Calls Out Overfitting in Agent Evolution
Google, Google DeepMind, UMD and UVA propose RRSI: don't lock down what agents can change about themselves — regularize how they search. Unregularized evolution hit 92.8 evolve score but only 40.3 OOD; RRSI reached 90.5/43.6 with fewer tokens. The data-backed case that self-improvement itself needs regularization.

The hottest direction in agent circles right now is called RSI — Recursive Self-Improvement. The idea is seductive: don't touch the model weights; let the agent rewrite "everything outside the model" — how prompts are written, how tools are called, how context is managed, how failures are recovered from — iterating round after round, getting stronger each time.
But a joint paper from Google, Google DeepMind, the University of Maryland, and the University of Virginia just poured data-backed cold water on the party: the "getting stronger" you're seeing might just be "getting familiar with the test." The paper proposes RRSI (Regularized Recursive Self-Improvement of agent harnesses), with a one-line thesis — self-improvement itself needs regularization.
Dating note first: the Chinese media report on this research was published October 6, 2026 (the page's date stamp reads "10.06 08:03," explicitly credited to 机器之心/Jiqizhixin, verified firsthand by our scout); the paper itself is on arXiv as 2609.24972 — September 2026. Code is open at google-research/rrsi; project page at regularized-rsi.com.
Why RSI caught fire: from "tuning models" to "tuning harnesses"
Start with why RSI caught fire. 2026 agent development has an open secret: models keep getting stronger, but "everything outside the model" is where the gap opens up. Two teams on the same model — one agent flies, the other is dumb as a brick. The difference isn't the model; it's the harness: prompt engineering, tool orchestration, memory management, error recovery — "the self outside the model."
RSI's logic: if the harness decides success, let the agent optimize the harness itself. Evaluate each change against the same evolve tasks, keep the good, discard the bad, loop. It reads like evolutionary algorithms meets LLMs, and in theory should compound. Over the past six months, everyone from academia to indie hackers piled into this direction.
The paper's authors first affirm the direction's value — then pivot: your evaluation method has a fatal structural flaw.
"Getting familiar with the test": three sins of the evolve score
The core problem: Harness Evolution evaluates changes round after round against the same evolve tasks. Machine learning has a classic name for this — tuning repeatedly on the validation set doesn't produce generalization, it produces validation-set overfitting. The paper splits RSI overfitting into three kinds:
Sin one: benchmark-specific fitting. The agent's modifications increasingly fit the quirks of this task set rather than getting genuinely stronger. Like a student memorizing mock-exam answers — a fresh exam exposes them.
Sin two: evaluation noise chasing. LLM outputs are stochastic; a one-time score bump may be luck. An unregularized evolution pipeline keeps "lucky" changes as "real improvements," accumulating superstition round after round until the harness is stuffed with voodoo.
Sin three: complexity accumulation. Each evolution round favors "adding things" — one more check, one more retry layer, one more prompt paragraph. Each addition looks reasonable alone; accumulated, the harness becomes a bloated monster — exploding maintenance cost, doubled inference tokens, collapse on any new scenario.
The paper delivers a counterintuitive gut punch of a number: on Agentic Workspace, Unregularized Evolution pushed the evolve score to 92.8 — above RRSI's 90.5. Higher evolution score. But on OOD (out-of-distribution tasks — genuinely unseen new tasks), the unregularized run scored only 40.3 versus RRSI's 43.6. The higher-scoring one was actually worse. That is the mathematical shape of "getting familiar with the test": every point gained on the evolve score bricks the OOD collapse.
And the cost stings more: per-trial Policy Token consumption was 3.80M unregularized versus 2.42M for RRSI. Unregularized evolution isn't just worse — it's more expensive, because it burns budget chasing noise.
RRSI's method: don't lock evolution, govern "how to search"
RRSI's cleverness is refusing the extreme. It doesn't say "stop self-evolving," nor "freeze the modification space." It says: evolution continues, but "how to search" must be regularized.
Concretely, two ends. The Proposal end (deciding what to try): not every modification deserves a trial — proposals pass screening first, so wild ideas don't waste evaluation budget. The Selection end (deciding which modifications earn permanence): a score bump alone doesn't qualify — stricter "permanence" criteria apply, including smoke tests and screening. The paper discloses that across 30 evolution rounds, only 10 candidates were truly accepted: 9 rejected for insufficient gains, 6 failed screening, 2 failed smoke tests.
That "10 out of 30" deserves a second look. It means that in an unregularized world, a good share of the 20 rejected changes were "looks like improvement, actually harmful." Accepted wholesale, they'd have quietly steered the harness off course. RRSI is essentially a quality inspector bolted onto evolution — better to reject a good change than admit a bad one.
In plain language: old RSI was "any study method stays if this exam's score went up"; RRSI is "a high score isn't enough — the study method must survive a different exam." The philosophy runs deep: it reapplies machine learning's oldest wisdom (regularization, cross-validation, Occam's razor) to the new species called agents.
A 30-year-old idea: RSI's past life
RSI isn't a 2026 coinage — it's a 30-plus-year-old idea that only recently turned from philosophy into engineering.
The concept traces to Jürgen Schmidhuber's 1987 dissertation: a machine that rewrites its own code, each rewrite making it smarter. Eliezer Yudkowsky made it central to AI safety discourse in the 2000s — "intelligence explosion": a self-improving AI entering a positive-feedback loop of ever-growing intelligence, eventually far beyond humans. Back then it was all philosophy and science fiction, because nobody knew how to implement "rewriting yourself."
The inflection came in 2023–2024. Large models crossed a capability threshold: they could write "decent" code, including code that improves themselves. The earliest shapes were "self-refine" tricks: have the model critique its own output, find flaws, rewrite, loop a few times — quality genuinely improved. Then came automated prompt optimization: frameworks like DSPy turned "tuning prompts" into a searchable optimization problem.
In 2025, RSI evolved from "tuning prompts" to "tuning harnesses." The landmark was papers proving that letting LLMs rewrite agents' tool-calling policies and memory management beat human experts' hand-tuning. From then on, "evolve tasks + looped evaluation" became the standard paradigm, and RSI moved from lab to engineering practice.
RRSI in 2026 is the latest link: when everyone's doing RSI, someone stops to ask "is your evaluation sound?" These "second-generation questions" — not "can we do it" but "are we doing it right" — usually mark a field's passage from hype to maturity. Just as deep learning around 2015 got serious about overfitting and generalization, RSI in 2026 getting serious about regularization means it's genuinely mainstream.
The historical read: RSI won't be a flash in the pan. From a 1987 vision to 2026 engineering took 39 years. RRSI isn't cold water on RSI — it's paving its road, turning "looks beautiful" into "actually reliable."
What exactly is a harness: unpacked for vibe developers
The paper's recurring "agent harness" may feel abstract to vibe developers. Unpacked, it's "everything outside the model" — five layers:
Layer one: prompts. How system prompts are written, how few-shot examples are chosen, how output formats are constrained. The shallowest, most editable layer — most "agent tuning" happens here.
Layer two: tools. Which tools the agent can call, how tool parameters are designed, how tool errors are handled. Tool design sets the agent's capability ceiling — hand an agent a dull knife and no amount of cleverness cuts.
Layer three: context/memory. How conversation history is compressed, how long-term memory is stored and retrieved, when retrieval triggers. The most hotly contested layer of 2026 agent engineering, with memory frameworks multiplying.
Layer four: control flow. How tasks decompose, how subtasks schedule, how many retries on failure, when to ask a human. The harness's "operating system" — it sets the agent's behavioral patterns.
Layer five: evaluation and recovery. How task completion is judged, how mistakes roll back, how abnormal states recover. This is where the paper's screening and smoke tests live — and the most neglected layer.
RSI's "self-improvement" modifies these five layers. RRSI's warning: the more you modify, the greater the overfitting risk — because each layer has its own "score-gaming" space. The prompt layer can memorize evolve-task answer patterns; the tool layer can specialize parameters to test tasks; the memory layer can remember test-set quirks. Five layers gaming scores together — how could the evolve score not look great?
The actionable advice: give your harness layered health checks. Every so often, ask five questions: does my prompt hardcode anything task-specific? Were my tool parameters tuned only on test tasks? Is my memory full of test-data features? Is my retry logic masking real bugs? When was my eval set last refreshed? If the answers make you uneasy, your harness may already be "gaming."
The data: validation across three domains
The paper didn't celebrate on a single benchmark — it validated across three domains:
Coding: evolved on Terminal-Bench 2.1, SWE-bench Verified rose from 82.0 to 83.8. Note the detail — evolution and validation used different benchmarks, which is OOD thinking embodied.
Agentic Workspace: JobBench 36.0 → 40.7, GDPval 48.8 → 52.3.
Engineering Design: Frontier-Eng 17.7 → 22.0.
Best OOD gain: +4.7 points. Not a dramatic number, but directionally consistent — all three domains positive. That's far more credible than "one benchmark up 20 points," which is usually overfitting itself.
The limitation, stated fairly: OOD validation spans 8 benchmarks. The "out-of-distribution" of 8 benchmarks and the "out-of-distribution" of real engineering tasks are separated by a chasm — real tasks have fuzzy requirements, legacy codebases, cross-team politics, none of which benchmarks simulate. The paper's methodological contribution is solid, but "does it transfer to real engineering" keeps its question mark — which the authors don't claim to have resolved either.
If RRSI is right: what agent infrastructure becomes
A thought experiment: if RRSI's thesis — "self-improvement needs regularization" — is widely accepted, what does agent infrastructure become?
First, evaluation infrastructure becomes its own business. Today's agent eval is "run benchmarks, read scores." Post-RRSI, eval must answer three questions: what's the evolve score, what's the OOD score, and how big is the gap (bigger gap = worse overfitting). Startups already sell "eval as a service"; RRSI hands them a new pitch: not just your score, but "is your evolution gaming the test."
Second, harness versioning becomes as important as code versioning. The paper's 30-rounds-10-accepted accept/reject log is itself a valuable asset. Future agent platforms may ship "harness git": every evolution a commit, with evolve score, OOD score, and token cost attached, rollback anytime. We covered prompt versioning for vibe projects before (cycle12); RRSI extends the need to the whole harness.
Third, "evolution budgets" become a new ops metric. The paper's disclosure — 3.80M tokens per trial unregularized vs 2.42M for RRSI, a 36% saving from regularization — turns real when agents self-evolve 24/7. CIOs will ask: how many tokens does your agent burn monthly on "self-improvement," and what's the ROI? Methods like RRSI go from "works better" to "costs less," and commercial adoption gets easier.
Finally, the deepest implication: RRSI hints at a new road for "agent alignment." Today's AI safety discourse mostly lives at model level (RLHF, constitutional AI). But future agent alignment may live at harness level: which is more controllable — an agent evolving unconstrained, or one whose evolution is regularized? RRSI's Proposal/Selection machinery is essentially brake pads on an agent's self-modification. Generalized, agent safety shifts from "governing models" to "governing the evolution process."
All speculation, of course. RRSI is one paper; 8 benchmarks is early validation. But a good paper's value was never just its numbers — it's the question it asks. And this paper's question — "is your agent really getting stronger, or just getting better at the test?" — deserves a spot on every agent builder's desk, glanced at regularly.
The dispute: who this paper challenges
RRSI's real gunpowder isn't in the data — it's in the methodological challenge. It takes on two deep-rooted assumptions in current agent engineering practice.
Assumption one: "higher training/evolution scores are always better." An intuition inherited from the deep-learning era — but the paper's 92.8-vs-90.5 reversal proves that in RSI, a higher score can mean worse overfitting. If that holds, multitudes of agent teams must rewrite their eval systems: you can't pop champagne just because the evolve curve climbs; you must watch the OOD curve too.
Assumption two: "the more autonomous the agent, the better." RSI's romantic vision is fully autonomous evolution, no humans needed. But RRSI's Proposal/Selection machinery essentially encodes human priors (what's worth trying, what counts as passing) as regularization terms. Fully autonomous evolution is inefficient; disciplined evolution is efficient — same as raising kids: disciplined autonomy beats total free-range.
Counter-voices exist. One argues RRSI's regularization is really "using 8 benchmarks' OOD to prevent overfitting," but those 8 benchmarks could themselves be fitted by future evolution — a "meta-overfitting" problem. Regularization just postpones overfitting a layer; it doesn't eliminate it. Fair criticism — but the paper's implicit reply is practical: postponing a layer is still progress; perfect generalization never existed.
A more pragmatic objection comes from engineering: 30 rounds accepting 10 candidates — who pays that "quality control" cost? Every screening and smoke test burns tokens and eval runs. For indie developers, RRSI's pipeline may be too heavy — you don't have Google's compute. The paper doesn't answer directly, but it open-sourced the code (google-research/rrsi), handing the "too heavy?" judgment to the community.
Takeaways for vibe/agent developers: the harness is the next battlefield
Finally, the practical meaning for our readers. Three judgments:
One: the harness is the next battlefield; models were the last one. In 2026, model gaps are narrowing (GPT-6, Claude, Gemini, and open models are hard to separate on most tasks) — real differentiation lives in harnesses. Methods like RSI/RRSI that automate harness optimization will become standard issue, like hyperparameter search before them: from black magic to infrastructure.
Two: give your agent "evolution discipline." Even without RRSI's full pipeline, its core ideas transfer free: hold out a set of tasks your evolution never sees; require smoke tests before any change merges to main; periodically purge harness logic "nobody knows why it exists." All three cost nothing. Do them tonight.
Three: beware "score illusion." Tuning an agent's prompt or workflow? Remember 92.8 vs 40.3: a rising score on your tuning set doesn't mean improvement. Build the habit: after every "optimization," test on a few completely unseen tasks. Score up but OOD down means overfitting — roll back.
In one sentence: RRSI's value isn't any specific algorithm — it's nailing an old saying back onto the agent era's wall: "there's no free lunch, and self-improvement is no exception." Evolution needs discipline; discipline needs regularization. Next time an agent brags "100 rounds of self-evolution, score up 30%," ask first: did you test OOD?
(Dating note: the 机器之心 report was published October 6, 2026, its page date stamp verified firsthand; the paper's arXiv ID 2609.24972 dates it to September 2026. Code and project links appear in the body.)
Sources
Related articles

Meta's SWE-sweep hides 4,068 real bugs across 100 repos in 22 languages — and gives agents no issue descriptions. The best setup fixes 75.3% with a human bug report, 4.8% without. The cliff shows that finding the problem, not writing the fix, is the human moat.

Microsoft's October 7 'hybrid intelligence' play: Microsoft Execution Containers GA on Windows 11 with policies enforced outside agent control; MAI-Code-1.1-Flash (137B/6.8B) quantized to 3-bit for devices; GitHub HydraFusion extended to Windows, routing tasks between local and cloud models. We break down the trio, the cost controversy, and what it means for indie developers.

In 2026, APIs are increasingly called by agents, not humans. This guide dissects the five pillars of agent-friendly API design — contracts, error codes, idempotency keys, pagination, and machine credentials — through real cases from Stripe and GitHub, plus a ready-to-use checklist, OpenAPI quality scorecard, and self-test prompt template.