Same architecture, different training
Mellum2 Thinking, which JetBrains open-sourced in June, was good at writing functions and completions but could barely work inside a repository on its own, running commands and fixing bugs: JetBrains measured just 2.0% on SWE-bench Verified. Mellum2.1-12B-A2.5B-Thinking, released around October 7, keeps the architecture: a mixture-of-experts model with 12 billion total parameters and only 2.5 billion active per token, with a 131,072-token context. Nearly all the change is in post-training: reinforcement learning in real software environments, teaching the model to work in a repository, run commands and call tools.
The result: SWE-bench Verified rises to 47.0%, SWE-bench Pro from 0 to 28.0%, and Terminal-Bench 2.1 from 0.6% to 17.4%. On non-agentic tests it scores 82.0% on LiveCodeBench v6 and 91.5% on HumanEval+.
Next to models its size
JetBrains re-ran the comparison models through its own pipeline, so the numbers don't always match their official cards. By its table:
- It clearly beats Gemma 4 E4B, for example 47.0% vs. 23.0% on SWE-bench Verified;
- Against Qwen3.5 9B it's mixed: ahead on LiveCodeBench and the BFCL v4 tool-calling test, but behind on SWE-bench Verified (47.0% vs. 50.0%), SWE-bench Pro (28.0% vs. 38.0%) and GPQA Diamond (64.6% vs. 77.8%).
In other words, it is strong at writing code, but as an agent fixing repositories and on general knowledge, it hasn't yet passed small models in its class such as Qwen3.5 9B.
Who should install it
With only 2.5 billion active parameters, inference cost and speed are close to a very small model, which suits it as a local or on-premises worker for coding agents: split-off subtasks, batch refactors, running tests and fixing small errors. The Apache 2.0 license allows commercial use. Besides BF16 weights, JetBrains published official GGUF quantizations such as Q4_K_M, Q6_K and Q8_0 that run directly in llama.cpp, Ollama and LM Studio; the multi-token prediction head for speculative decoding is still "coming soon," and the model card lists no hosting provider yet. All scores are JetBrains' own, and the agentic results depend on the specific harness it used.
via: Hugging Face model card, JetBrains blog, MarkTechPost report