Back to Explore
NewsVibeFix 编辑部Updated Oct 11, 2026

TianxiCode Tops SWE-bench-Live: Lenovo's Code Agent Hits 71% Solve Rate with DeepSeek-v4.1-Flash

Lenovo Tianxi AI's in-house code-agent framework TianxiCode, paired with DeepSeek-v4.1-Flash, topped the SWE-bench-Live Lite leaderboard at a 71% solve rate with official Verified certification. We unpack why this "real engineering" benchmark is harder, what "framework > model" really means, and three takeaways for vibe coding practitioners.

Conceptual illustration of TianxiCode ranking first on the SWE-bench-Live leaderboard with a 71% solve rate

71% to the Top: A Homegrown Agent Framework Reaches the Summit for the First Time

On October 8, the official SWE-bench-Live leaderboard quietly updated. For the first time, the top spot on the Lite split went to a Chinese name: TianxiCode — a professional code-agent framework developed in-house by Lenovo Tianxi AI. It topped the world with a 71% problem-resolution rate: 213 out of 300 real-world engineering tasks solved, edging out the second-place entry at 70.33% by less than a percentage point. But first place is first place. More importantly, this summit came with a seal: the leaderboard stamped TianxiCode with Verified certification — it passed the official full review, and the score is real and valid.

SWE-bench-Live Lite 榜单对比(示意图)

What makes this story interesting is not another "China takes No. 1" headline, but the combination of three keywords: in-house framework, homegrown open-source model, real-world engineering benchmark. The paired model is DeepSeek-v4.1-Flash — Chinese, open-source, and lightweight, not any closed-source giant's flagship. A "lightweight engine" paired with a self-developed agent engineering framework outperformed every rival on what is widely considered the hardest code benchmark in the world. That is exactly the judgment the coverage keeps hammering home: framework > model.

The news was first reported by QbitAI on October 9 and subsequently corroborated by English-language media: Lite split, 300 tasks, 213 solved, 71% to the top, leaderboard entry dated October 8. The numbers line up across sources — this is a hard news story that holds up to verification.

Why SWE-bench-Live Is Harder Than SWE-bench

To grasp what this No. 1 is worth, you first need to understand what kind of exam SWE-bench-Live is. The classic SWE-bench became the "gaokao" of code-model evaluation the moment it launched in 2024: tasks drawn from real GitHub repository issues, requiring models to generate patches that pass tests. But it had a fatal flaw: the task set was static. Once the dataset went public, models might have "seen" those issues during training — data contamination, in industry parlance. Scores kept climbing, but nobody could say whether that reflected real skill or memorized answers.

SWE-bench-Live was built to fix exactly that. Its task pool is not published all at once; instead, it continuously harvests real issues from the newest, most actively maintained GitHub projects, updating on a rolling basis. Every task is a "fresh question" for contestants: no memorized answers, no peeking at test cases. What gets tested is an agent's full ability to locate a problem from scratch in a real codebase, understand context, and write a patch. Each of the 300 tasks on the Lite split comes from a pit some real developer actually fell into.

The Verified certification is this exam's second line of defense. SWE-bench-Live's official verification mechanism requires competing teams to submit full agent run trajectories — from the first second the issue is read, every retrieval, every tool call, every trial-and-error step must leave a trace. Official reviewers check each one: was there any answer leakage? Any peeking at hidden test cases? Is the trajectory complete and reproducible? Only teams that pass this audit cleanly earn the Verified mark next to their score.

In other words, TianxiCode's 71% was not "gamed." It was earned on a continuously updating task pool, with the entire problem-solving process handed over and audited by the officials. That is why the industry treats SWE-bench-Live as the closest thing to a real measure of "engineering capability" for code agents — it does not test whether a model can write code, but whether an agent system can independently fix a bug in a real project.

The trajectory-review design itself deserves a second look. It amounts to the officials saying: we don't just check whether your patch is correct, we check how you got there — whether you took shortcuts, whether you "happened" to guess the hidden tests' intent. This kind of "process audit" is far stricter than "result scoring," and much closer to how code review works in real engineering: a patch nobody can understand how it was written should never be merged, no matter how many tests it passes. A team willing to hand over its full trajectories for review has already won a point on engineering honesty.

A Footnote on the Numbers: 71% vs. 70.33%

First place at 71%, second at 70.33% — a gap of just 0.67 percentage points, roughly two tasks out of 300. A margin this thin says something important: among the top players, model capability no longer separates anyone. Everyone is running first-tier models; the deciding factor is not "whose model is smarter" but "whose engineering framework better converts model smarts into patches." That is the "framework > model" thesis, confirmed by the leaderboard itself.

There is also a detail that often gets overlooked: 71% means 87 tasks went unsolved. Even the world champion fails nearly three in ten real-world engineering problems. That number is both a ceiling and an honest signal — a reminder that code agents still have a long way to go before "fully reliable," and that between topping a leaderboard and being production-ready lies the long tail of that 29%.

"Framework > Model": What That Claim Is Really Worth

One line in the QbitAI coverage jumps off the page: even calling a top-tier closed-source model, without top-level agent architecture orchestrating it, often leaves you helpless against real engineering problems. It reads like vendor PR, but the leaderboard data backs it up — the winner was not the biggest model, but an engineering framework paired with a lightweight open-source one.

This judgment echoes a broader cognitive shift sweeping the agent track in 2026: models are becoming commodities. The code-capability gap between flagship models is narrowing while prices fall fast. What truly separates contenders is the layer above the model — the agent framework: how it retrieves, how it plans, how it manages context, how it validates its own output. Models set the floor; frameworks set the ceiling.

TianxiCode's three officially disclosed capabilities translate, in plain language, into the three hurdles every code agent must clear:

  • Cross-file multi-hop retrieval + precise context management: real bugs never live in a single file. The root cause of a null pointer can hide three call layers away, and a fix may require understanding conventions across five or six files at once. An agent's retrieval ability determines how far it can "see" into a codebase, while context management determines whether it loses critical information along a long reasoning chain. This is the first life-or-death line for every code agent.
  • Self-driven planning + multi-round tool use: a good agent is not a "question-and-answer" parrot but an actor that decomposes tasks itself, decides which tool to use next, and adjusts its plan based on what comes back. Across the multi-round loop of "locate → reproduce → fix → verify," every decision depends on planning ability. If planning collapses, more tool calls are just spinning wheels.
  • Closed-loop patch generation + test-driven self-healing: the most underrated link. Generating a patch is easy; the hard part is running tests on it yourself afterward, reading the failures, rolling back, and trying again — forming a "write → test → fix" self-healing loop. An agent without this loop is essentially a fancy autocomplete; with it, it starts to resemble an "engineer" that can stand behind its own results.

My take: of the three, the third is the dividing line for code agents in 2026. The first two answer "can it do the work"; the third answers "can the work be trusted." SWE-bench-Live only recognizes the third — a patch counts only if it passes real tests, no matter how pretty the process looked. So the number 71% is really saying: TianxiCode's self-healing loop held up across 213 real bugs.

Looked at from the other side, "framework > model" carries a deeper meaning: it drags the center of competition away from "burning money to train ever-bigger models" and back to "doing solid engineering." That is good news for teams without a hundred-billion-parameter training budget — innovation happens at the architecture layer, and architecture-layer innovation runs on understanding of real problems, not compute.

TianxiCode 三大能力(示意图)

What Vibe Coding Readers Should Take Away: What Makes a Good Agent

Back to our readers' daily lives: for those of us writing code every day with Cursor, Claude Code, or Codex, what is actually worth taking from this story? Three practical takeaways.

First, stop obsessing over "which model" you switched to. Many people's agent-workflow optimization stalls at "waiting for the next stronger model." TianxiCode's summit proves there is still enormous headroom in framework engineering at current model levels. Instead of chasing new models, spend the effort on prompt engineering, context pruning, toolchain orchestration — the "framework layer" stuff. They offer far bigger leverage on success rates. A well-tuned workflow on an older model routinely beats a raw new model running naked.

Second, the test loop is the ticket from vibe-coding "toy" to "production." Getting an agent to write code is easy; getting it to stand behind what it wrote is hard. One practical habit: always wire up a "run tests → read failures → auto-fix" loop for your agent instead of serving as its human tester. TianxiCode's "test-driven self-healing" sounds lofty, but at the personal-workflow level it reduces to one simple rule — never let the agent say "done" before the tests go green. The harder your acceptance bar, the harder its output.

Third, lightweight model + strong framework may be the best value-for-money combo. DeepSeek-v4.1-Flash is a lightweight, cheap, open-source model, and its ability to top the board shows: with a strong enough framework, you don't need to pay flagship prices for every call. For individual developers and one-person companies, that means agent-workflow costs can drop dramatically — and the saved budget buys better retrieval, longer context, and more verification rounds, which pay back more. The smart way to spend is on "letting the agent try more times," not on "using the priciest model every time."

A Glimpse of China's Agent Infrastructure Landscape

Zoom out, and TianxiCode's summit is not an isolated event. In the same week, Huawei-backed openJiuwen open-sourced its enterprise-grade AgentOS, packaging multi-agent orchestration, self-evolution, and sandbox gateways into a complete Apache 2.0 offering. Meanwhile, Lenovo has stated that TianxiCode will mature into the code-intelligence capability inside its AI hardware products. Open-source infrastructure on one side, on-device hardware deployment on the other — Chinese agents are moving from the "publish papers, top leaderboards" phase into the "build the foundation, ship the product" phase.

Lenovo Tianxi's plan is especially worth savoring. A PC maker building a code agent looks like a stretch until you think it through: the AI PC narrative has been running for two years, and the market is still waiting for a "killer" on-device agent scenario. Code intelligence happens to be one of the most mature agent scenarios with the strongest willingness to pay. If TianxiCode's capabilities make it into Lenovo's AI hardware lineup — say, a local coding assistant on a developer laptop — then it stops being a leaderboard score and becomes a real product line.

Of course, a sober note is in order. SWE-bench-Live measures one dimension: "fixing bugs." Real engineering also involves understanding requirements, designing architectures, and debugging across systems — things no leaderboard measures. The 71% summit deserves applause, but what it proves is "this framework is the best in the world at fixing bugs," not "AI can replace engineers." Several real-world hurdles still stand between the two. The best attitude toward leaderboard results: neither dismiss their engineering substance nor hallucinate them into a victory of general capability.

One more angle: Lenovo Tianxi's choice to build its own framework rather than "wrapping" models is itself a statement of direction. While many teams are still competing over "how many models we've plugged in," Tianxi put the work into the agent-architecture layer. It's a slower, heavier road — but once traveled, what accumulates is something nobody can take away: understanding of engineering details, accumulation of failure modes, design experience with closed loops. The No. 1 spot is simply the first time the outside world has seen that accumulation.

Epilogue: The Summit Is Only a Starting Point

Looking back, the most thought-provoking part of this story: the winner was not the biggest model, but an engineering framework that pushed a lightweight model to its limits. In 2026, as model capabilities rapidly commoditize, that may be a more important signal than "China's No. 1" — the agent era's competition is shifting from "whose model is smarter" to "whose engineering is more solid." Investment in frameworks, obsession with closed loops, ruthlessness about verification — this "unglamorous work" is becoming the widest moat.

TianxiCode taking 71% together with Verified sets a high bar for China's agent infrastructure: win on audits, not marketing; win on engineering, not parameters. What to watch next is whether these capabilities actually make it into Lenovo's hardware products and become something developers use every day. Leaderboard champions come every year; champions that ship in products are rare.

Views 0Comments 0

Comments (0)

ME
0/1000
Loading comments...

Sources

Browse projectsPublish your project

Related articles

DHH on stage at Rails World announcing 37signals has stopped writing code by hand
News
"We're Done Writing Code by Hand": Rails Creator DHH's Agent-Era Manifesto

Rails creator DHH announced at Rails World that 37signals is 'done writing code by hand.' This piece unpacks the November 2025 inflection point, the move to native apps and Rust, the counter-evidence from the same newsletter — and what vibe coders should actually take away.

AI CodingIndustry TrendsProduct News
Illustration of the Nemotron dual-gold recipe: SFT and RL checkpoints, 22000 curated programming problems, and the GenCorrect generate-evaluate-refine inference loop
News
NVIDIA Open-Sources Its 'Double Gold' Training Recipe: Nemotron Beats Top Human IOI Score, Full 22,000-Problem Dataset Released

NVIDIA's Nemotron systems hit gold level at IOI 2026 (535.4/600, above the top human score of 498.27 in an unofficial run) and IMO 2026 (30/42, graded by official IMO graders) — and the team open-sourced the full recipe: SFT/RL checkpoints, both training datasets, a new 200-problem olympiad benchmark, inference pipelines, and prompts. The lesson is co-design of model, data, and inference loop: GenCorrect's generate-evaluate-refine cycle carried a 291-point model past the 438.3 gold bar.

Open-source ProjectsModel UpdatesAI Coding
OpenAI and Ironclad partner to train GPT-6 Astra on 11 scored contracting tasks: 55.0% vs 41.6% and 19.2 vs 37.0 minutes against GPT-5.6 Sol
News
OpenAI Finds Its Agents a Sparring Partner: Ironclad's Contracting Workflows Become 11 Training Tasks, Astra Beats Sol by 32%

OpenAI's blog 'Advancing computer use with Ironclad' marks a paradigm shift: computer use moves from general capability to per-application customized RL training. GPT-6 Astra, the first frontier model trained on Ironclad tasks, scored 55.0% vs 41.6% for GPT-5.6 Sol on 11 contracting tasks (8-50 scoring criteria each), with time per attempt falling from 37.0 to 19.2 minutes. The moat is moving from models to the partner list - and OpenAI is openly recruiting the next batch of software companies.

Product NewsIndustry TrendsAI Coding