Back to Explore
NewsVibeFix 编辑部Updated Oct 9, 2026

JetBrains Open-Sources Mellum 2.1: 12B MoE Reasoning Model Activating Only 2.5B Parameters per Token, Built to Work for Coding Agents

JetBrains has open-sourced Mellum 2.1, a 12B MoE model activating only 2.5B parameters per token. With its architecture unchanged since June, reinforcement learning in real environments lifted SWE-bench Verified from 2.0 to 47.0. Apache 2.0 licensed and self-hostable, it is positioned as the fast, cheap execution layer for coding agents — strong on coding and tool use, still trailing Qwen3.5-9B on the hardest agentic tasks.

Mellum 2.1 benchmark comparison chart against Mellum 2, Qwen3.5-9B and Gemma 4 E4B, published on the JetBrains AI blog
JetBrains AI 博客发布的 Mellum 2.1 评测对比图:与 Mellum 2、Qwen3.5-9B、Gemma 4 E4B 的基准测试成绩

A model that trained itself into a "hired hand"

JetBrains has open-sourced Mellum 2.1. The company famous for IntelliJ IDEA and PyCharm is doing something few IDE vendors attempt: training its own models, open-sourcing them, and aiming them squarely at coding agents as worker models. One note on timing: the official blog post carries no explicit date stamp; its URL path reads 2026/10. Multiple independent outlets — MarkTechPost, oossa, tech-insider — consistently date it October 8, 2026, which is the date we use.

Mellum 2.1's parameter count won't scare anyone: a 12B mixture-of-experts model activating 2.5B parameters per token, released under Apache 2.0 and already on Hugging Face. What's worth discussing is its positioning — JetBrains states it plainly on the official blog: it's built for coding agents and "fast sub-agents that run on your own hardware." In other words, Mellum 2.1 doesn't aspire to be an omniscient brain; it wants to be the cheap, fast, always-on executor inside agentic systems.

That positioning is remarkably clear-headed amid 2026's coding-agent frenzy. While everyone else competes on whose giant model's SWE-bench score is higher, JetBrains went the other way: it kept the Mellum 2 architecture from June completely unchanged and spent the entire summer on one thing — post-training. The result: SWE-bench Verified jumped from 2.0 to 47.0, more than a 23× leap. Same architecture, transformed capability. That story deserves a close read.

Under the hood: MoE architecture unchanged; the "work skills" are new

First, the hard specs. Mellum 2.1 is a 12B-total-parameter mixture-of-experts model: 64 experts, 8 active per token, 2.5B parameters actually participating in computation. That's the essence of MoE — inference cost scales with active parameters, not total ones. A 2.5B activation footprint means its per-token inference cost is roughly that of a 2.5B dense small model, while the 64 experts provide far greater "knowledge capacity." This is the mainstream survival strategy for small models in 2026: the Qwen and Gemma families' same-class open models are all going sparse.

The context window is 131K tokens — adequate for an agent worker: reading a few source files, reviewing test output, doing a round of edit-and-verify rarely needs million-token context. Notably, JetBrains also prepared a multi-token prediction (MTP) head for Mellum 2.1, for speculative decoding in vLLM. Official figures claim MTP delivers roughly 1.6× speedup in single-request scenarios. The MTP head is currently marked "coming soon," but the direction is clear: push latency down to levels acceptable for high-frequency agent calls.

Beyond architecture, every bit of improvement comes from post-training — RL (reinforcement learning)-led post-training. JetBrains uses a refreshingly honest formulation on the blog: RL went from "a short final stage" to "the main part of training." Three things happened in practice.

First, RL tasks were massively expanded, spanning math, competitive programming, science, tool use, and software engineering. The data mix combines open RL datasets with self-built tasks. JetBrains' attitude toward open data is old-school and correct: filter before training — because open data is riddled with "broken tests, unverifiable answers, or tasks that are too easy or impossible for the model." That sentence reads bland, but anyone who has done RL training knows data-cleaning quality directly sets the RL ceiling.

Second, JetBrains built thousands of RL environments in-house and launched millions of sandboxed runs over the course of training. This is the key to Mellum 2.1's transformation from "code completer" to "agent worker": the model no longer just predicts the next token on static text — it acts in real environments, exploring codebases, editing files, running tests, checking its own changes, and learning from environmental feedback. "Real-environment RL" is standard practice for agent-model training in 2026; DeepSeek, Kimi, and others all use it. JetBrains simply applied it to a small model with 2.5B active parameters.

Third, and most thought-provoking: JetBrains says plainly that they "ran many experiments on both the training methods and the data, and kept what held up." No mythology, no "we discovered a magic trick" — just honest experimental iteration. That reads as credible precisely because it is: post-training gains are earned one controlled experiment at a time.

Benchmarks head-to-head: strengths are real, weaknesses are on the record

JetBrains ran this evaluation with unusual decency: the same pipeline, the same evaluation setup, comparing Mellum 2.1 against Mellum 2 (Thinking variant) and two same-class open models — Qwen3.5-9B and Gemma 4 E4B. Agentic benchmarks used an open-source agent harness (Pi v0.73.1, shell and file tools), 114K context, up to 16K tokens per turn, with Mellum 2.1 sampled at temperature 1.0. Methodology this transparent is rare in vendor self-evaluations. One caveat, though: JetBrains ran the benchmarks itself, and MarkTechPost notes that "vendors' own cards report different numbers" for Qwen3.5-9B and Gemma 4 E4B — keep a grain of salt handy.

Start with agentic coding, where Mellum 2.1 improved most. SWE-bench Verified leapt from Mellum 2's 2.0 to 47.0. Read that number against two yardsticks: vertically, a 23× jump proves RL post-training genuinely works for small models on agent tasks; horizontally, Qwen3.5-9B scores 50.0 and Gemma 4 E4B 23.0 under the same pipeline. Mellum 2.1 has essentially caught up with Qwen3.5-9B on SWE-bench Verified — but hasn't surpassed it.

Now the harder tests. On SWE-bench Pro, Mellum 2.1 scores 28.0 versus Qwen3.5-9B's 38.0 — a 10-point gap. On Terminal-Bench 2.1, Mellum 2.1 gets 17.4 against Qwen3.5-9B's 21.7. JetBrains didn't hide this gap on its blog, and we won't airbrush it either. Terminal-Bench tests real task completion in terminal environments: installing dependencies, invoking commands, handling messy errors. A model that hits 47 on SWE-bench but only 17 on Terminal-Bench is fluent at "fixing bugs along an established workflow" but still shaky at "finding its own way through a chaotic real terminal." That is precisely the capability an agent worker needs most — and Mellum 2.1's most honest weakness.

BenchmarkMellum 2.1Mellum 2Qwen3.5-9BGemma 4 E4B
SWE-bench Verified47.02.050.023.0
SWE-bench Pro28.00.038.04.0
Terminal-Bench 2.117.40.621.73.4
LiveCodeBench v682.069.475.469.4
BFCL v4 (tool use)62.349.658.552.5
AIME 25/26 (math)83.360.186.745.0
GPQA Diamond (science)64.651.077.853.1

Pure coding ability is Mellum 2.1's brightest calling card. LiveCodeBench v6 at 82.0 beats Qwen3.5-9B's 75.4 and Gemma 4 E4B's 69.4 — best in class. HumanEval+ 91.5 and MBPP+ 79.4 also lead the group. Tool use (BFCL v4) at 62.3 is likewise best in class — and for an agent worker, tool-calling ability arguably matters more than raw code-writing, since it determines whether the model faithfully follows the agent harness's instructions to invoke shells and read/write files. On math, AIME 25/26 hits 83.3, close to Qwen3.5-9B's 86.7; on science QA (GPQA Diamond), 64.6 trails Qwen's 77.8 noticeably, suggesting the RL dividends concentrate in code and toolchains rather than general scientific reasoning.

Reading the report card horizontally: Mellum 2.1 is a "lopsided but usefully lopsided" model. Writing code, calling tools, fixing bugs along a defined workflow — the highest-frequency agent subtasks — it's best in class; but on wild terminal-environment tasks and ultra-hard long-horizon agent work (SWE-bench Pro), Qwen3.5-9B still wins. That profile fits its positioning exactly: it was never meant to be the omniscient brain; it's here to be the execution layer.

A methodology footnote: this evaluation's transparency deserves credit

A word on evaluation methodology, because it affects how much to trust these numbers. JetBrains disclosed details vendors usually keep quiet: all four models under one pipeline, one agent harness (Pi v0.73.1, shell plus file tools), 114K context, up to 16K tokens per turn, even Mellum 2.1's sampling temperature (1.0). Mellum 2 was re-evaluated under the new pipeline — its numbers differ slightly from the June technical report, and the company said so upfront. Putting re-evaluation discrepancies on the table like that is rare in vendor self-testing.

But see the other side of the coin: every number is self-reported by JetBrains — the Hugging Face model card states "All values are self-reported by JetBrains" verbatim. MarkTechPost added its own footnote when republishing the comparison table: vendors' own cards may report different numbers for Qwen3.5-9B and Gemma 4 E4B. So the correct reading: treat this as a relative ranking under one consistent yardstick, not absolute truth. The real test comes when independent evaluators (like BenchLM) re-run with their own harnesses. So far it's a good sign: BenchLM's SWE-bench Verified leaderboard already lists Mellum2.1-12B-A2.5B-Thinking at 47%, matching the official 47.0.

Speed: nearly 2× Qwen3.5-9B under heavy load

Speed is what lets Mellum 2.1 credibly call itself a "fast sub-agent." Official figures: on a single H200 under heavy load, Mellum 2.1's token throughput is nearly double Qwen3.5-9B's; in single-request scenarios, MTP adds roughly 1.6× acceleration. The gap comes from architecture: MoE activates only 2.5B parameters per token, while Qwen3.5-9B activates 9B densely (its official architecture description is a hybrid gated DeltaNet plus gated attention — actual activation far above 2.5B). The activation gap translates directly into an inference-FLOPs gap, and then into throughput and cost gaps.

For agent systems, speed isn't just "nice UX" — it's "usable at all." One agent task can invoke the model dozens or hundreds of times: reading files, grepping, attempting edits, running tests, reading errors, editing again. If every call queues behind a cloud flagship model, the agent's wall-clock time gets wrecked. A locally deployed small model with double the throughput compresses that high-frequency call latency into acceptable territory. That's why JetBrains keeps hammering "on your own hardware" — speed and deployment location are bound together.

Positioning: why "hired hand" is a good business

Mellum 2.1's official use cases list three items: a capable worker inside agentic systems, a general assistant, and private self-hosted deployment. The first is the soul. JetBrains' vision: inside a coding agent, Mellum 2.1 can handle "different parts of an agent's plan" — from identifying the root cause of a failing test to drafting and checking a fix. Translation: big models decompose tasks and make decisions; small models do the dirty work.

This division of labor is no longer novel in 2026. Anthropic's sub-agent mechanism and the "planner-executor" pattern across agent frameworks all tell the same story: not every call deserves a flagship model. A senior engineer doesn't personally run every unit test suite — CI does that. Same for agents: high-frequency, low-difficulty, verifiable subtasks (reading code, editing files, running tests, checking results) go to cheap, fast small models; genuine judgment — planning, architectural decisions — stays with cloud giants.

The economics are blunt. Cloud flagship APIs charge per token, and a complex agent task burning several dollars is routine; a 2.5B-active model running locally has near-zero marginal cost (electricity and hardware depreciation aside). For individual developers and small teams, that means "agent freedom": previously you hesitated to let agents run loose for fear of the bill; now, swapping the execution layer to a self-hosted small model collapses the cost of experimentation. JetBrains made this point back at Mellum 2's June launch: it aimed to solve production AI's three mountains — latency, throughput, cost. Version 2.1 lands that argument in the agent scenario.

Of course, the division only works if the small model is "good enough." When Mellum 2 launched, SWE-bench Verified sat at 2.0 — at that level a worker only creates extra work, with the agent spending more tokens correcting its mistakes than it saves. Version 2.1 pushing that number to 47.0 is exactly the point: it has crossed the "hired hand that doesn't get in the way" threshold for the first time. 47 isn't 90 — it will still err on complex tasks — but inside an agent system with a verification loop (edit, run tests, retry on failure), a worker at this level already generates net positive value.

Open-source ecosystem: Apache 2.0 + self-hosting aims at compliance demand

Mellum 2.1 ships under Apache 2.0: commercial use, modification, and redistribution allowed, no royalties. For open models, the license choice is itself a positioning statement — JetBrains wants "take it to production," not "research demo." Compare: some open-weight models use community licenses restricting commercial use, which makes enterprise legal teams flinch; Apache 2.0 is essentially a pre-cleared pass for enterprise procurement.

More important is the "private, self-hosted deployment" use case. JetBrains lists it separately on the blog: run Mellum 2.1 locally or on your own infrastructure, keeping code and data fully under your control. Who is that for? Finance, government, healthcare, large enterprises — places where code may never leave the intranet and cloud APIs, however strong, can't get in. Those customers previously could only watch AI coding from the sidelines or pay dearly for privatized deployments. Now a 12B MoE under Apache 2.0 sits on Hugging Face, with a Q4_K_M quantization at 8.1GB (a third-party GGUF repository already lists five quantization builds, Q4_K_M recommended) — the self-hosting bar has been pressed down to "one workstation with a decent GPU."

Distribution today: Hugging Face is live; GGUF builds (llama.cpp / Ollama / LM Studio) and the vLLM MTP speculative-decoding head are marked coming soon. Interestingly, MarkTechPost reports a third-party GGUF repository already hosting five quantization files — the community's hands are faster than the official release. That's usually an early signal of a well-received open model: when the quantization community bothers to adapt it on day one, people genuinely want to run it, not just watch.

One layer deeper: JetBrains is securing its own AI supply chain. Tech-insider's commentary nailed it: JetBrains is betting that "a small, fast, openly licensed model can do more for coding agents than another oversized proprietary one." If every autocomplete and every agent call inside its IDEs rents intelligence from OpenAI, Anthropic, or Google, JetBrains' AI strategy stays hostage forever. From June's Mellum 2 to October's 2.1, JetBrains is walking the "own models, own IDEs, own agent stack" path. For JetBrains AI users, IDE agent calls will likely default to in-house models in the future — cost, latency, and data all under its own control.

Practical value for vibe coders: run the numbers on local agents first

For our readers — vibe coders who write code with AI daily — Mellum 2.1's value isn't "yet another open model." It's that it could change your agent-running cost structure. Let's run the numbers.

Hardware bar first. The 12B MoE's BF16 full weights run about 24.3GB; Q8_0 quantization 12.9GB; Q4_K_M 8.1GB (third-party GGUF repository figures, Q4_K_M recommended, 88.0% top-token match). What does 8.1GB mean? A 16GB consumer GPU (RTX 4080 class) running Q4_K_M: weights take just over 8GB, the rest goes to KV cache — full 131K context won't fit, but real agent-worker scenarios (tens of K of context) do. Mac users needn't fret either: the MLX community typically follows GGUF releases, and Apple Silicon's unified memory handles 8B-class quantized models comfortably. Official Ollama / LM Studio GGUF builds are still "coming soon" — the impatient can watch third-party quantization repos; the cautious should wait for official builds.

Now the usage bill. Say you do daily development with agent tools like Claude Code or Pi: one medium-complexity task (fix a bug, add a feature module) routinely burns hundreds of thousands of tokens. At flagship cloud prices, that's several dollars; a dozen tasks a day adds up to hundreds of dollars a month in API bills without trying. If you point your agent's execution layer at a local Mellum 2.1 endpoint — planning and hard decisions still on the cloud giant, high-frequency subtasks (reading files, editing code, running tests, verifying) local — the token bill gets slashed. Hardware is a one-time investment, electricity a fixed cost, marginal call cost near zero. That "hybrid routing" is exactly the usage JetBrains described at Mellum 2's June launch: routing, sub-agents, private AI.

Three practical tiers. Tinkering: pull JetBrains/Mellum2.1-12B-A2.5B-Thinking from Hugging Face, run it under vLLM or transformers, and test it on your own repos — official benchmarks are nice, but nothing beats localizing one failing test in your own codebase. Waiting: hold for official GGUF on Ollama/LM Studio — one command and you're running, the lowest-friction path for vibe coders. Production: if you're building internal coding agents for a team or working in compliance-sensitive industries, the technical groundwork can start now — the license is clean and the self-hosting story is complete, which is more than many "open-weight but awkwardly licensed" models offer.

One caution: Mellum 2.1's weaknesses (Terminal-Bench 17.4, SWE-bench Pro 28.0) mean don't expect it to solo complex tasks. Use it as the execution layer with a "tests are truth" verification loop — edits must run tests, failures auto-retry — and its 47-point SWE-bench level converts into real productivity. A small-model agent without a verification loop is just a confident bug factory; that holds for every small model.

Competitive landscape: Qwen still holds the crown, but it's a different arena

The landscape in one sentence: on hardcore agent tasks, Qwen3.5-9B remains the same-class open-model champion (SWE-bench Pro 38.0 and Terminal-Bench 21.7 both best in group); but Mellum 2.1 opened a different arena — nearly 2× the speed, 2.5B active parameters per token, Apache 2.0 for unrestricted commercial use — playing the "cheap, fast agent execution layer" card. Gemma 4 E4B looks more like an also-ran in this comparison, trailing on most fronts. The real contest isn't "whose benchmark is highest" but "who becomes the default local worker in agent systems" — and that seat is still empty.

Closing: post-training ate the architecture gap

What Mellum 2.1 deserves to be remembered for isn't any single score but one fact: architecture completely unchanged, SWE-bench Verified from 2.0 to 47.0. That suggests post-training matters more for agent capability than we used to think. A small model with 2.5B active parameters, via real-environment RL and millions of sandboxed runs, can nip at the heels of 9B-class dense models — a gentle blow to the "stack more parameters" faith, and a strong endorsement of the "train small models well" path.

For the open-source ecosystem, Mellum 2.1 fills a critical puzzle piece: flagships in the cloud for brains, and now a cheap, fast, cleanly licensed, self-hostable pair of "hands" on the open side. The brain-and-hands division of agent systems finally has respectable hands in the open world. JetBrains' next move is worth watching: if it deeply integrates the Mellum line into its IDEs' agent flows, JetBrains AI could become the first IDE agent stack that "runs its own open small model by default" — a landmark step in IDE vendors evolving from "AI feature integrators" into "AI supply-chain players."

For vibe coders, the test can be beautifully simple: once the official GGUF lands, pull it down and run one real task on your own project. Whether it's any good, your test suite will tell you — more honest than any benchmark.

Sources

Browse projectsPublish your project

Related articles

Concept illustration of Anthropic OSS Scanner: AI models scanning open-source code for security vulnerabilities
News
Anthropic Launches OSS Scanner: Free AI Security Scans for Open Source, Raw Model Reports Straight to Maintainers

Anthropic's free opt-in OSS Scanner uses its strongest models, including Claude Mythos, to periodically scan open-source projects. Reports are fully model-generated with no human review. Models surfaced 29,000+ candidate vulnerabilities in six months; PostgreSQL, OpenSSL, wolfSSL, and curl all responded positively. We break down the mechanics, the data, and the controversies.

Product LaunchSecurity & PrivacyOpen-source Projects
Developer reviewing code with an AI reasoning model in a code editor
News
Vercel AI Gateway Ships Stealth Reasoning Model Glyph Cluster, Free During Stealth: A Hands-On Testing Guide and the Data-Privacy Red Line

Vercel added stealth reasoning model Glyph Cluster to AI Gateway (Oct 7): coding plus long-context analysis, function calling, streaming. Pro/Enterprise teams with purchased credits use it free during stealth (stealth/glyph-cluster), selectable in Claude Code, Codex, Cursor. Why indie developers should benchmark it, the gateway's failover and budget value, and the limits: text-only input, no structured outputs, no ZDR — prompts may train the model, so set a data boundary first.

Model UpdatesProduct LaunchAI Coding
REA concept illustration: a coding agent analyzing a binary through MCP decompilation tools
News
REA Gains 13K Stars in a Day to Top GitHub Trending: A Decompiler for Your Coding Agent

On October 9, REA (Reverse Engineer Anything) gained roughly 13,000 GitHub stars in a single day, topping GitHub Trending's dev-tools daily ranking. It wraps decompilation and static analysis into an MCP server + CLI - one npx rea-agents setup lets 12 coding agents, from Claude Code to Cursor, read binaries with no source code. We break down its methodology, capability map, and the gray areas of reverse engineering.

Product NewsDeveloper WorkflowAI Coding