← AUTOMATER NEWSROOM

JetBrains Releases Mellum2.1, an Apache-Licensed Model for Self-Hosted Coding Agents

Mellum2.1 keeps its 12B mixture-of-experts architecture while expanding reinforcement learning for repository work. JetBrains reports coding gains, but deployment artifacts and speed claims need separate checks.

A conceptual explore, edit and check loop with reinforcement-learning feedback, beneath Mellum2.1’s unchanged 12B total and 2.5B active parameter architecture.A conceptual explore, edit and check loop with reinforcement-learning feedback, beneath Mellum2.1’s unchanged 12B total and 2.5B active parameter architecture.
Original explanatory graphic based on JetBrains’ Mellum2.1 announcement. The loop illustrates repository-oriented reinforcement learning, not a measured benchmark or exact training implementation.

JetBrains has released Mellum2.1, an Apache 2.0-licensed model aimed at coding agents and smaller workers running on users’ own infrastructure. The company says the update improves repository exploration, file editing and checking changes through expanded reinforcement learning, while retaining Mellum2’s 12-billion-parameter mixture-of-experts architecture with 2.5 billion active parameters.

The consequential change is in training after pre-training. In its release announcement, JetBrains describes millions of sandboxed runs across thousands of environments. That makes this a release focused on learning to perform software work inside an environment, rather than an architectural expansion. The performance evidence presented in the supplied announcement remains JetBrains’ own evaluation.

JetBrains Blog’s Mellum2.1 release announcement page.
Source: JetBrains Blog. Publisher announcement introducing Mellum2.1 and its focus on coding agents trained in real environments. · Original source

What changed in training

JetBrains says reinforcement learning moved from a short final stage to the main part of training for this version. It added tasks covering software engineering, tool use, mathematics, competitive programming and science, combining open datasets with internally developed tasks.

The company also describes filtering those sources for broken tests, unverifiable answers and tasks that were either too easy or impossible for the model. This matters because a training environment’s feedback determines which behavior gets rewarded. A coding agent needs useful signals about whether an edit works, beyond whether its explanation sounds plausible. The announcement explains the direction of that work, but does not supply enough detail to independently assess the filtering process or reproduce its results.

JetBrains’ news archive also lists the release, describing Mellum2.1 as a model trained with reinforcement learning in real environments for coding agents and fast sub-agents. That corroborates the publisher’s positioning; it is another page from the same company, not an independent performance assessment.

The architecture stays compact, but total weights still matter

The earlier Mellum2 technical report, submitted on May 29, supplies architectural background. It describes 64 experts with eight active, grouped-query attention, sliding-window attention and a multi-token prediction head designed to support speculative decoding. JetBrains says Mellum2.1 retains the architecture of version 2.

The distinction between total and active parameters is useful for deployment decisions. Activating 2.5 billion parameters per token reduces the computation involved relative to activating all 12 billion. It does not mean the model has only 2.5 billion parameters to store. Hardware planning still depends on the full weights, precision, serving software, context length and workload.

The older report supports an explanation of the design. Its benchmark results should not be presented as measurements of Mellum2.1, whose central change is subsequent training.

Availability and speed are separate questions

The release announcement says Mellum2.1 is available on Hugging Face. It describes GGUF builds for llama.cpp, Ollama and LM Studio, along with the separate multi-token prediction head for speculative decoding in vLLM, as coming soon. Those additional artifacts should therefore remain rollout items unless their exact repositories and contents are verified separately.

JetBrains reports comparisons with Mellum2, Qwen3.5-9B and Gemma 4 E4B using the same evaluation setup. It says agentic coding improved most, with gains also appearing in coding, mathematics, tool calling and general knowledge. The supplied text does not include numerical benchmark scores, so it does not support a precise ranking of task success.

Its speed claims also describe different operating conditions: almost twice Qwen3.5-9B’s token throughput under heavy load, and about a 1.6-fold single-request acceleration with multi-token prediction. These are vendor-reported results, not guaranteed deployment speeds. The MTP claim also needs to be read alongside the announcement’s pending release of the separate head.

Why this release matters

Mellum2.1 offers a candidate for teams that want a locally operated worker to investigate failing tests, propose edits and check a fix. Its license and self-hosting focus make it relevant to organizations weighing control over code and infrastructure against the operational work of serving a model.

The useful next comparison is completed repository work: success rate, elapsed time, retries and infrastructure cost under a consistent agent harness. Our guide to pricing self-hosted runs per completed task provides the relevant evaluation frame. Faster token generation can help, but an agent’s value ultimately depends on whether it finishes the assigned job correctly.

JetBrains has announced a concrete post-training update to an existing open model. The remaining questions are how those gains transfer to other repositories and harnesses, and when the promised deployment artifacts become verifiably available.

Sources

  1. Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents - The JetBrains Blog
  2. News Archives - The JetBrains Blog
  3. [2605.31268] Mellum2 Technical Report