JetBrains Open-Sources Mellum 2.1 Thinking: A 12B Mixture-of-Experts With 2.5B Active, SWE-bench Verified Up From 2% to 47%, Still Slightly Behind Qwen3.5 9B

Around October 7 JetBrains released Mellum2.1-12B-A2.5B-Thinking on Hugging Face under Apache 2.0: a mixture-of-experts reasoning model with 12 billion total parameters and 2.5 billion active per token, a 131,072-token context, built for coding agents that run commands and call tools inside repositories and for private self-hosted deployment. The architecture is the same as June's Mellum2 Thinking; the gains come mainly from reinforcement learning post-training in real software environments. JetBrains' own SWE-bench Verified score rises from 2.0% to 47.0%, with 82.0% on LiveCodeBench v6; in the same table Qwen3.5 9B scores 50.0% on SWE-bench Verified and 38.0% on SWE-bench Pro (Mellum 2.1: 28.0%). Official GGUF quantizations are already available for llama.cpp, Ollama and similar runtimes; a multi-token prediction head for speculative decoding is listed as coming soon.

Same architecture, different training

Mellum2 Thinking, which JetBrains open-sourced in June, was good at writing functions and completions but could barely work inside a repository on its own, running commands and fixing bugs: JetBrains measured just 2.0% on SWE-bench Verified. Mellum2.1-12B-A2.5B-Thinking, released around October 7, keeps the architecture: a mixture-of-experts model with 12 billion total parameters and only 2.5 billion active per token, with a 131,072-token context. Nearly all the change is in post-training: reinforcement learning in real software environments, teaching the model to work in a repository, run commands and call tools.

The result: SWE-bench Verified rises to 47.0%, SWE-bench Pro from 0 to 28.0%, and Terminal-Bench 2.1 from 0.6% to 17.4%. On non-agentic tests it scores 82.0% on LiveCodeBench v6 and 91.5% on HumanEval+.

Next to models its size

JetBrains re-ran the comparison models through its own pipeline, so the numbers don't always match their official cards. By its table:

  • It clearly beats Gemma 4 E4B, for example 47.0% vs. 23.0% on SWE-bench Verified;
  • Against Qwen3.5 9B it's mixed: ahead on LiveCodeBench and the BFCL v4 tool-calling test, but behind on SWE-bench Verified (47.0% vs. 50.0%), SWE-bench Pro (28.0% vs. 38.0%) and GPQA Diamond (64.6% vs. 77.8%).

In other words, it is strong at writing code, but as an agent fixing repositories and on general knowledge, it hasn't yet passed small models in its class such as Qwen3.5 9B.

Who should install it

With only 2.5 billion active parameters, inference cost and speed are close to a very small model, which suits it as a local or on-premises worker for coding agents: split-off subtasks, batch refactors, running tests and fixing small errors. The Apache 2.0 license allows commercial use. Besides BF16 weights, JetBrains published official GGUF quantizations such as Q4_K_M, Q6_K and Q8_0 that run directly in llama.cpp, Ollama and LM Studio; the multi-token prediction head for speculative decoding is still "coming soon," and the model card lists no hosting provider yet. All scores are JetBrains' own, and the agentic results depend on the specific harness it used.

via: Hugging Face model card, JetBrains blog, MarkTechPost report