Back to Explore
NewsVibeFix 编辑部Updated Oct 10, 2026

75% With a Bug Report, 5% Without: Meta's SWE-sweep Exposes AI Coding Agents' Biggest Blind Spot

Meta's SWE-sweep hides 4,068 real bugs across 100 repos in 22 languages — and gives agents no issue descriptions. The best setup fixes 75.3% with a human bug report, 4.8% without. The cliff shows that finding the problem, not writing the fix, is the human moat.

Developer debugging code in an IDE, illustrating the SWE-sweep benchmark for autonomous bug discovery

The same top configuration, tested on the same 20 repositories: hand it a human-written bug report and it fixes 75.3% of the bugs. Take the report away, change nothing else, and it fixes 4.8%. A 70-percentage-point cliff, on the same model and the same agent framework — the only variable is whether someone first told it what was broken, and how.

Those numbers come from SWE-sweep, a new AI coding-agent benchmark published by Meta researchers on October 1. Its pitch is blunt: how many bugs can LMs find & fix in large codebases? The benchmark reportedly spans 100 open-source repositories, 22 languages, and 4,068 real bugs — but its most aggressive design choice isn't scale, it's a subtraction: no issue descriptions.

The agent gets a repository frozen at a fixed commit and one broad instruction: "find and fix as many bugs as you can." No issue titles, no reproduction steps, no filenames or line numbers. It has to explore the repo, read the docs, run the tests, and decide for itself what looks wrong — before it ever gets to fixing. That first step is exactly the one every previous coding benchmark did for the agent.

The timing is telling. In 2026's vibe-coding wave, the "AI programmer" narrative has slid from "helps you write code" to "does the work for you" — auto-fixing bugs, nightly repo sweeps, unsupervised refactors; pitch decks promise it all. Yet nearly every score propping up that narrative comes from benchmarks that test solving. SWE-sweep reads like Meta pouring a bucket of data-flavored cold water over the industry: before we talk about replacement, let's test whether your agent can take the first step unescorted.

First, the record straight: this is not SWE-bench

This site covered the SWE-bench Pro October leaderboard on October 6, so readers may conflate the two — the distinction matters. SWE-bench (Pro included) tests solving: you get the codebase plus a user-written issue explaining what broke and roughly how to reproduce it. The agent just has to fix it. It measures whether the fix is right and how fast it lands.

SWE-sweep tests discovery: the issue itself has been escorted out of the exam room. The agent must first answer "is anything broken here?" and only then "how do I fix it?" To use an exam analogy: SWE-bench is an open-book test where the teacher also highlighted the key chapters; SWE-sweep drops you into a 400,000-line codebase and says "there are dozens of mistakes in here — find them yourself." They measure fundamentally different abilities, and the scores are not comparable — which is why SWE-sweep's leader sits at 4.7% while SWE-bench leaderboards routinely show 70%+.

This distinction determines how you read the numbers: 4.7% doesn't mean the models got dumber. It means the exam got harder — harder in that the examiners finally started testing the question everyone had been skipping.

Look one level deeper and SWE-bench's high scores start to look inflated. A well-written issue has already done the two hardest steps of debugging for the agent: localization ("it's probably the payments module") and characterization ("expected behavior is X, actual is Y"). What remains is often just translating a natural-language description into a code diff — precisely what large language models are best at. So part of the "exponential rise in AI coding ability" we've watched over the past two years may actually be a rise in humans describing problems more clearly, not in the models themselves. SWE-sweep's value is that it weighs the two separately for the first time. From now on, when you see a model's coding score, your first reflex should be to check the test conditions: how many hints did the issue give? The more hints, the more water in the score.

How 4,068 bugs got "hidden" in there

SWE-sweep's construction deserves a close look, because it determines how much to trust the benchmark. The researchers collected real issue–pull-request pairs from open-source repos, then for each repository identified the commit in history where the most bugs coexisted, freezing that moment with all its concurrent bugs in place. Every bug ships with two artifacts: a fix patch and a test patch.

At evaluation time, the agent's modified codebase faces two gates. First, the repository's original test suite is restored and run in full to confirm no regressions — the site states it plainly: the agent is explicitly told not to break existing tests, and doing so zeroes the entire run. Second, each bug gets a set of hidden tests, at least one of which fails on the unmodified codebase (fail-to-pass) and all of which must pass after the fix.

Debugging code: SWE-sweep requires agents to autonomously discover and fix bugs in large codebases

Then comes the crucial filter: only bugs "discoverable from the repository alone" are admitted. "Discoverable" means the expected behavior must be inferable from inside the repo — documentation, type annotations, existing tests, caller code, invariants, naming standards, or an unambiguously bad smell like a crash or data loss. All but 2 of the benchmark's bugs have concrete in-repo contracts describing expected behavior. That closes an obvious loophole: if even a human would need a god's-eye view to spot the bug, testing an agent on it would be a rigged game.

The site even publishes the full prompt given to the agent, and reading it you can feel the designers' anti-cheating obsession. It defines what counts as an acceptable source of expected behavior: (1) user-facing documentation, including docstrings, type annotations, and external standards a module claims to implement; (2) universally shared expectations — no segfaults in C++, no accidental deletion of user data, no frozen UIs. "Out-of-band knowledge" is categorically excluded: what the maintainers intended, what the original issue text said, what the model remembers about the project from training — none of it counts. The agent is also required to run the test suite at the start and memorize which tests were already red: only those pre-existing failures may stay red; touching anything else zeroes the run. In other words, this benchmark tests not just "can you find the bug" but "are your hands steady when you operate."

One more note: every model on the leaderboard runs the same agent scaffold, mini-SWE-agent. The site's FAQ preempts the obvious question — wouldn't a better scaffold score higher? The paper ran ablations with Claude Code and Codex, and neither significantly outperformed mini-SWE-agent; if anything, mini-SWE-agent beat Codex by a clear margin. So this leaderboard mostly measures gaps between the models themselves.

The leaderboard: honestly low scores, frighteningly high bills

The live leaderboard on the official site (at the time of writing):

  • #1: Sol 5.6 (xhigh) — 4.7%, $7,230 per full run
  • #2: Luna 5.6 (xhigh) — 2.5%, $224 per full run
  • #3: Terra 5.6 (xhigh) — 1.5%, $357 per full run
  • #4: Luna 5.6 (high) — 1.4%, $28 per full run
  • #5: Opus 5 (xhigh, Anthropic) — 1.3%, $5,363 per full run
  • #6: Kimi K3 (Moonshot AI) — 0.6%, $2,451 per full run

Three comparisons are worth chewing on. First, the leader Sol 5.6 paid $7,230 for its 4.7% — over seven grand burned on a single full sweep — and the site adds that even higher reasoning modes are being tried, which could push a single run past $10,000. The current recipe for "finding bugs" is brute-forcing long trajectories with money, and the marginal returns diminish: the site states explicitly that repeated attempts recover additional bugs but with diminishing gains — and that extending a trajectory can even undo earlier repairs.

Second, the real value pick is #2, Luna 5.6 (xhigh): 2.5% is more than half the leader's score at $224 — roughly 1/32 of the cost. If you wanted to use this benchmark as a "nightly repo sweep" tool, the Luna 5.6 tier is the realistic choice, even though 2.5% in absolute terms is still bleak.

Third, Anthropic's Opus 5 spent $5,363 for 1.3%, and Kimi K3 spent $2,451 for 0.6%. Money and score are completely uncorrelated here — which suggests "autonomous bug discovery" isn't something you can buy your way up linearly. There are real capability gaps between models.

Engineer working at night with AI data streams: autonomous bug discovery still demands heavy compute

One easily missed detail: since every model runs the same mini-SWE-agent scaffold — and the FAQ confirms ablations with Claude Code and Codex didn't meaningfully beat it — the ceiling on this leaderboard is essentially the ceiling of the "model + generic scaffold" combination, not some buried engineering trick. Breaking past 5% will likely require a new agent paradigm, not longer trajectories — and the official word on longer trajectories is already in: diminishing returns, with a risk of breaking what you already fixed.

The takeaway: after the cliff, the human's position is clearer than ever

The 75.3%-to-4.8% cliff doesn't expose the models so much as it exposes our measurement illusion about AI coding over the past two years. SWE-bench-style benchmarks have been measuring "how fast the model swallows once a human chews the problem for it," and we misread those scores as "how independently the model can work." SWE-sweep simply removed the hand that had been holding the model up — and the scores showed their true shape.

What does this mean for the vibe-coding era? It means "who finds the problem" is the human moat, not "who writes the fix." Writing the fix is something models already do well: 75.3% with a report in hand is three out of four, and then some. But "noticing something is off" draws a near-blank. That maps exactly onto the most valuable and least outsourceable part of real development: the instinct when reading an error trace, the nose for "this logic smells wrong," the ability to translate a user's vague "it feels a bit slow" into a reproducible problem.

The second narrative this deflates is "unsupervised automatic bug fixing." If you're planning to let an agent sweep your repo every night and collect PRs in the morning, print this leaderboard and tape it to your monitor: the strongest configuration, no hints, $7,230 per run, under 5% fixed. This isn't a "wait for stronger models" problem — in the discovery dimension, the current agent paradigm (long-horizon exploration plus trial and error) is inherently expensive and inefficient. The unsupervised bug hunter, as of October 2026, is a myth.

There's a second-level reading of "75% vs 5%," too: a bug report is, in essence, a human doing both "discovery" and "localization" for the agent in one shot. A good report's "expected vs actual behavior" defines the problem; its reproduction steps fence in the scope. What remains — "write a patch that turns the test green" — is exactly the pattern-matching models excel at. So the two ends of the cliff test nearly two different kinds of intelligence: the 5% end asks you to "find the leak in a building with no signage," the 75% end asks "someone's pointing at the leak — can you fix the pipe?" The former needs a world model and active exploration; the latter needs code generation — one is the models' weakness, the other their strength. That split carries more information than any debate over "AI developer replacement" timelines.

One blunt closing thought: the most stinging thing about this benchmark isn't the score — it's that it redefines what "programming ability" is made of. We used to assume "writes code = can program," so the faster models wrote code, the more anxious we got. Once SWE-sweep splits programming into "finding problems" and "solving problems," the anxiety points somewhere new — not "AI writes code faster than me" but "can I spot problems earlier and more accurately than AI?" The good news: that's exactly where experience, business understanding, and user empathy pay off most — and the hardest part to quantify on any benchmark. The moat was always there. We were just measuring in the wrong place.

Three practical conclusions for indie developers

1. Always pair the agent with a signal source before letting it work. The leaderboard proves agents are excellent fix executors and terrible scouts. In practice: a failing test, an error trace, a user report, a crash log — any concrete signal drags the agent from the 4.8% hell back up to the 75% heaven. Workflow-wise, "humans localize to the file level, agents fix" remains the highest-ROI posture in 2026.

2. Treat "writing bug reports" as a core skill. It sounds counterintuitive: in the AI era, one of the most important skills is writing issues? But the data is the data — a report that states "what the expected behavior is, what the actual behavior is, how to reproduce it" makes the same model 15x more productive. A person who writes good reports commands agents an order of magnitude more effectively than one who doesn't. For indie developers that's good news: this skill costs no money, only care.

3. Don't pay a premium for "autonomous discovery." If you run agents via API for code review or scheduled scans, watch the cost structure: the xhigh reasoning tier costs dozens of times more per run than the high tier, for a gain of a few tenths of a point. The saner strategy right now: use cheap tiers for broad suspicious-spot flagging (as a linter), and keep the "is this actually a bug?" judgment with the human.

4. Steal SWE-sweep's prompt trick: make the agent run the test baseline before touching anything. The most copy-worthy move in the official prompt is forcing the agent to run the full suite first and memorize what was already red. The same applies when you let agents edit code day to day: establish a "what the world looked like before my change" baseline, and the diff instantly shows whether it fixed or wrecked things. Most agent-induced disasters trace back to a missing baseline — was the regression introduced by the agent, or was it already broken? Without a baseline you can't tell. With one, every agent edit is auditable, which is the precondition for letting it near production code at all.

Being honest: this benchmark has its own boundaries

A proper report shouldn't only relay the official strengths. SWE-sweep's own FAQ concedes two boundaries worth surfacing. First, it is still a history-derived benchmark: every bug comes from an issue–PR pair a human already found, fixed, and merged. So it measures "of the bugs humans could find, how many can agents find" — bugs humans never found aren't in the denominator at all. In the site's own words, history-derived benchmarks necessarily under-count the bugs actually present in a repo. So 4.7% can be read optimistically as "a lower bound" or pessimistically as "only 4.7% of even the human-found subset" — both readings hold.

Second, the agent is denied internet access and may not look at any code outside the repo. A real developer hunting a bug would Google, check Stack Overflow, read the upstream changelog — "out-of-band knowledge" counts as cheating in the benchmark and as basic competence in real life. So SWE-sweep measures "pure read-the-repo" discovery ability: a conservative, lab-conditions number. Read it as a lower-bound estimate of autonomous discovery, not as the agent's true level.

Neither weakens the conclusion. If anything: a conservative test still produced a 70-point cliff, which means the "discovery" gap isn't an artifact of harsh conditions but a real capability chasm. And there's more to look forward to: the site says public submissions are coming, and more models are being evaluated (the FAQ names Astra, Fable, and 5.5 as not-yet-listed). As new models take their turn on this board, we'll have — for the first time — a ruler for measuring "discovery ability." The next time you see "our agent fixes bugs autonomously" marketing, you'll know the first question to ask: what's your SWE-sweep score?

Dating note and sources

For transparency on the date: the swesweep.com pages carry no explicit date stamp. This article dates the release to October 1, 2026, based on three-way corroboration — the official repository facebookresearch/swe-sweep was created at 2026-10-01T04:59:29Z (verified firsthand via the GitHub API); VibeLeaderboard's intel entry is labeled "Published Oct 1, 2026, source swesweep.com"; and generativeai.pub's deep dive confirms "Meta researchers published on October 1." All three agree.

Primary sources: the SWE-sweep site (full leaderboard and evaluation details) and the facebookresearch/swe-sweep repository. Leaderboard figures and evaluation methodology are quoted from the official site; the 75.3%-vs-4.8% controlled comparison (same 20 repositories) is cited from generativeai.pub's analysis.

Sources

Browse projectsPublish your project

Related articles

Mellum 2.1 benchmark comparison chart against Mellum 2, Qwen3.5-9B and Gemma 4 E4B, published on the JetBrains AI blog
News
JetBrains Open-Sources Mellum 2.1: 12B MoE Reasoning Model Activating Only 2.5B Parameters per Token, Built to Work for Coding Agents

JetBrains has open-sourced Mellum 2.1, a 12B MoE model activating only 2.5B parameters per token. With its architecture unchanged since June, reinforcement learning in real environments lifted SWE-bench Verified from 2.0 to 47.0. Apache 2.0 licensed and self-hostable, it is positioned as the fast, cheap execution layer for coding agents — strong on coding and tool use, still trailing Qwen3.5-9B on the hardest agentic tasks.

Product LaunchAI CodingModel Updates
A developer writing code at a computer, with API documentation and a terminal window on screen
Guide
Make Your Product Callable by Agents: An Agent-Friendly API Design Guide

In 2026, APIs are increasingly called by agents, not humans. This guide dissects the five pillars of agent-friendly API design — contracts, error codes, idempotency keys, pagination, and machine credentials — through real cases from Stripe and GitHub, plus a ready-to-use checklist, OpenAPI quality scorecard, and self-test prompt template.

Backend EngineeringAI CodingDeveloper Workflow