Dev News Daily ENDE

JetBrains ships Mellum2.1: same 12B open model, retrained for coding agents

JetBrains released Mellum2.1 on 8 October, the next version of the open model it published in June. The architecture is unchanged — a 12B mixture-of-experts model with 2.5B active parameters, under the Apache 2.0 licence — and JetBrains says everything that changed happened after pre-training.

What changed. Mellum2 was fast but, in JetBrains' words, could not work inside a repository at the level it wanted. For 2.1 the company made reinforcement learning the main part of training rather than a short last stage:

  • new RL tasks in maths, competitive programming, science, tool use and software engineering, mixing open datasets with tasks JetBrains built itself — and filtering every source, because open RL data often has broken tests, unverifiable answers or tasks that are trivial or impossible;
  • in-house infrastructure running thousands of environments, with millions of sandboxes launched over training, so the model learns to explore a codebase, edit files and check its own changes.
JetBrains ships Mellum2.1: same 12B open model, retrained for coding agents
JetBrains ships Mellum2.1: same 12B open model, retrained for coding agents — Dev News Daily

How it compares. JetBrains benchmarked it against Mellum2 and two open models of similar class, Qwen3.5-9B and Gemma 4 E4B, with one evaluation setup for all. It reports the largest gain in agentic coding and smaller gains in coding, competitive programming, maths, tool calling and general knowledge. On speed: because the architecture is the same, 2.1 is as fast as Mellum2; JetBrains says it is the fastest of the group under heavy load, serving almost twice as many tokens as Qwen3.5-9B, and that multi-token prediction makes a single request about 1.6 times faster. These are the vendor's own measurements.

Where it fits. The pitch is a worker inside an agent — finding the root cause of a failing test, drafting and checking a fix — and private deployment on your own hardware so code does not leave the building. The weights are on Hugging Face now; GGUF builds for llama.cpp, Ollama and LM Studio, and the MTP head for speculative decoding in vLLM, are announced as coming soon. Teams already running a local model for sub-agents can swap it in and measure on their own repositories, which is the only benchmark that settles the question.

Source: JetBrains AI Blog, "Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents", 8 October 2026 — https://blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/

Written by Victoria Shinder.