MiMo V2.6 is open weights from Xiaomi, and its 9B model runs at 24.09 tokens a second on a 48 GB Mac

Xiaomi released the MiMo V2.6 family as open weights on September 21, 2026. We ran the 9B model on a 48 GB Mac and give the maker's figures for the two large ones, which we have not run.

0:00

Xiaomi published the MiMo V2.6 family as open weights on Hugging Face on September 21, 2026, with an MIT licence tag on every repository. The largest model, MiMo-V2.6-Pro-RL, holds 1,024,216,603,392 parameters in 573.46 GB of weight files, according to the Hugging Face API. We ran the smallest one, a 9.4 billion parameter fine-tune of Qwen3.5-9B, on our test machine, an Apple M4 Pro with 48 GB of memory. An 8-bit GGUF build of it decoded at a median of 24.09 tokens a second over five runs. The two large models are 3.7 and 11.9 times the size of that machine’s memory, so we have not run them.

Takeaway points

  • Hugging Face counts 1,024,216,603,392 parameters in MiMo-V2.6-Pro-RL and 310,756,322,688 in MiMo-V2.6-Flash-RL.
  • The two large models are 573.46 GB and 177.74 GB of weight files, so neither fits the 48 GB of memory on our test machine.
  • We ran the 9B distilled model: it wrote at a median of 24.09 tokens a second over five runs, and its process held 10.80 GB of memory.

What is MiMo V2.6, and how big are the models?

MiMo V2.6 is a set of five repositories in the XiaomiMiMo organisation on Hugging Face. Three were created on September 21: MiMo-V2.6-Pro-RL and MiMo-V2.6-Flash-RL at 15:39 UTC and MiMo-V2.6-Distill-Qwen-9B at 18:18 UTC, the API shows. MiMo-V2.6-Pro-MOPD and MiMo-V2.6-Flash-MOPD followed on September 27.

Both large models are mixture-of-experts designs, in which each token passes through a few small sub-networks instead of the whole model. The Pro card lists 70 layers with 384 routed experts, 8 of them chosen per token, and gives the size as “1.02T total / 42B activated parameters”. The Flash card lists 48 layers and 256 experts, again with 8 active. The config files in the two repositories carry the same layer and expert counts and set the maximum context at 1,048,576 positions.

Xiaomi gives two totals for Flash. The model card says 309 billion parameters and the technical report says 310 billion. Hugging Face counts 310,756,322,688 in the weight files.

Repository Parameters (Hugging Face count) Weight files Active per token (card)
MiMo-V2.6-Pro-RL 1,024,216,603,392 132 files, 573.46 GB 42 billion
MiMo-V2.6-Flash-RL 310,756,322,688 67 files, 177.74 GB 15 billion
MiMo-V2.6-Distill-Qwen-9B 9,409,813,744 4 files, 18.82 GB not a mixture-of-experts model

By our arithmetic the two large repositories hold about 4.5 bits per parameter. The technical report says Xiaomi trained the models with the MXFP4 number format from the mid-training stage onward, and the config files name mxfp4 as the storage type. The API lists 1,000,190,509,056 of the Pro model’s parameters under an 8-bit integer type and 13,378,781,184 as 8-bit floating point.

The cards say the models take text, images, video and audio. Each adds a 681 million parameter vision encoder and two audio encoders of 308 million and 127 million parameters.

All five repositories are tagged MIT and none is gated. None of them holds a LICENSE file: a request for one in each repository returned a 404 on October 5. The Qwen3.5-9B repository that the small model is built on is tagged Apache 2.0.

What do RL and MOPD mean in MiMo V2.6 model names?

No card defines the RL suffix in so many words. The Pro-RL card describes the family as built to scale reinforcement learning, and the MOPD cards call the earlier upload the “RL-stage” checkpoint. The technical report says the reinforcement learning stage cost $2.6 million for Pro and $0.9 million for Flash. Each training step used 1,568 prompts with 16 attempts apiece, and coding made up 68% of the tasks, according to the report.

MOPD is defined. Xiaomi’s blog post of September 27 expands it as “Multi-teacher On-Policy Distillation”, a method in which several specialist teacher models supervise the released model’s own outputs. The cards call the version used here MOPD2 and expand that as “Multi-Prefix Multi-Teacher On-Policy Distillation”.

The MOPD checkpoints exist because of a fault. “Following the release of MiMo-V2.6, tool-call repetition emerged as one of the most noticeable issues affecting the user experience,” the blog post says. The model would issue the same request to a tool again and again, which the MOPD card describes as “appearing busy while making no progress”.

Xiaomi measured the fault by counting identical calls inside a single turn. In the OpenCode agent the rate was 1.02% for Flash-RL and 0.54% for Pro-RL, the post says, and in Claude Code it was 0.27% and 0.10%. Xiaomi calls its measure a “lower bound on observable repetition”, because it leaves out repeats across turns and calls that differ slightly.

The company traced the fault to its own training rule. A rollout was stopped and scored zero only when the model issued more than 32 tool calls in one turn, the post says, so heavy calling below that line went unpunished and grew. Retraining with a limit of eight would have cost an estimated $2.31 million. Xiaomi instead trained a specialist teacher and distilled it into both models for about $90,000. The post says repetition rates “dropped substantially” afterwards and gives the results as charts, not as a table.

The MOPD repositories report the same parameter counts as the RL ones, and their config files are byte-for-byte identical. Xiaomi left the RL checkpoints online. Its API has served the updated models since September 25 under unchanged names, the post says.

How do the MiMo V2.6 benchmark scores compare with Claude Opus 5 and GPT-5.6 Sol?

Xiaomi’s own table puts MiMo V2.6 Pro close to the two rival models on some tests and well behind on others. Every score on the cards is Xiaomi’s. The technical report says three of the test sets (MiMo Code Bench, MiMo Cyber Bench and MiMo Visual Coding) are internal, and that rival models were run at their highest reasoning setting. The report claims “performance comparable to that of frontier models across various domains”.

Benchmark (Xiaomi’s table) MiMo V2.6 Rival models
DeepSWE v1.1 Pro 71.9, Flash 67.9 Claude Opus 5 74.0, GPT-5.6 Sol 73.0
AutomationBench v1.0.6 Pro 53.1 Claude Opus 5 50.3
Terminal Bench 4.0 Pro 34.9 Claude Opus 5 49.0
ExploitBench Pro 47.9, Flash 25.3 GPT-5.6 Sol 78.5

On DeepSWE v1.1 the previous model, MiMo-V2.5 Pro, scored 19.0 in the same table. On AutomationBench the table has Pro ahead of Claude Opus 5. On Terminal Bench 4.0 and ExploitBench the gaps run the other way.

The report does not agree with itself on one benchmark. Its text says the DeepSWE score rose to 72.6 for Pro and 65.7 for Flash over the course of reinforcement learning. Its final table gives 71.9 and 67.9. The table’s columns read “MiMo-V2.6 Pro” and “MiMo-V2.6 Flash” with no suffix, so the report does not say which of the two published checkpoints each score belongs to. The MOPD cards carry no benchmark table.

We have not run any of these benchmarks. They are Xiaomi’s figures, and we repeat them only to show what the maker claims.

Does MiMo V2.6 Distill 9B run on a 48 GB Mac?

Yes. A Q8_0 GGUF build of MiMo V2.6 Distill 9B loaded and ran on our test machine at a median decode speed of 24.09 tokens a second across five runs, with every layer on the GPU.

The official weights of the two large models do not fit. The Pro repository’s 573.46 GB of weight files is about twelve times 48 GB, and the Flash repository’s 177.74 GB is more than three and a half times. The cards give no memory requirement. They give launch commands for two server engines, SGLang and vLLM, and the Pro command for SGLang is written for two nodes. No MiMo-V2.6 repository in Xiaomi’s organisation holds a GGUF or MLX build.

The small model is a different kind of thing. Its card says it was made “through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data”, 77.4 billion tokens of it. Its config sets a context of 262,144 positions, a quarter of the large models’ figure. The GGUF file we ran came from the ggml-org repository on Hugging Face, not from Xiaomi: MiMo-V2.6-Distill-Qwen-9B-Q8_0.gguf, 9,527,498,048 bytes. Its sha256 checksum matched the value Hugging Face lists for the file. That is about a fifth of 48 GB, and about half the size of the 18.82 GB of 16-bit weights in Xiaomi’s own repository.

We ran it through the llama.cpp Metal runtime inside LM Studio (llama.cpp-mac-arm64-apple-metal-advsimd 2.47.0) on macOS 27.0.1, on October 5, 2026. We loaded it with a context of 16,384 tokens. The file sat on an external USB drive, a T7 Shield formatted exFAT, and was read from there. Before the run we unloaded Gemma 4 31B QAT, the model from our test of Gemma 4 31B QAT on an M4 Pro with 48 GB, to free memory.

Each of the five runs used a different prompt, so the runtime’s prompt cache could not help. Every prompt was a block of shipment notes followed by a question on a different topic, 300 tokens were allowed, the temperature was 0.7 and the reply was streamed. The cache trap is explained in our guide to measuring tokens per second in LM Studio.

Run Prompt tokens Tokens generated Time to first token Prefill, tokens a second Decode, tokens a second
1 1,553 300 31.38 s 49.5 25.76
2 2,083 300 5.02 s 414.6 24.61
3 2,644 300 7.41 s 356.8 24.06
4 3,202 243 9.76 s 328.0 24.09
5 3,764 258 11.62 s 324.0 23.97

The median decode speed was 24.09 tokens a second, with a range of 23.97 to 25.76. The median prefill speed was 328.0 tokens a second, with a range of 49.5 to 414.6.

Run 1 is the odd one. It was the first read of the weights from the external drive, so its prefill figure of 49.5 tokens a second includes loading and says little about the model. Its decode speed, 25.76 tokens a second, was the highest of the five. Our reading is that the prefill median of 328.0 is a fair summary of runs 2 to 5, which ran from 324.0 to 414.6, and that run 1 should be set aside for prefill.

After the five runs the llama-server process held 10.80 GB of resident memory, from the ps reading. Our reading is that 10.80 GB against a 9.53 GB file is consistent with the weights plus the 16,384 token context. We did not measure the split.

The swap readings need care. Before the model was loaded, with Gemma still loaded, swap used was 23,389.25 MB of 24,576 MB, with 5.7 GB of free and inactive memory. After Gemma was unloaded and before MiMo was loaded, free and inactive memory was 28.9 GB and swap used was 8,479.06 MB of 24,576 MB. After the five runs, swap used was 8,271.94 MB of a total that had shrunk to 9,216 MB. The machine was already swapping before the test began. We cannot say from these readings how much swap, if any, the model caused.

Have we run MiMo V2.6, and what has Xiaomi not published?

We have run one MiMo V2.6 model: the 9B distilled one, as the GGUF build above, five times, on one machine, on one day. Five runs at prompt lengths of 1,553 to 3,764 tokens do not tell you the speed at other context lengths. We did not test answer quality, so we make no claim about whether the model’s replies are good. We did not load the vision file, mmproj-MiMo-V2.6-Distill-Qwen-9B-Q8_0.gguf, which we downloaded. We tried no other runtime. For a different model, Ollama vs LM Studio on an M4 Pro compared two runtimes on the same file, and we have not done that for MiMo.

We have not run MiMo-V2.6 Pro or Flash, and we have no figure for either. The Pro and Flash weight files are 573.46 GB and 177.74 GB, and neither fits in our test machine’s memory.

Xiaomi has not published a licence file in any of the five repositories. The technical report says the company is also releasing its training environments and a training framework, and we found no link to either in the report or on the cards. The report does not name the checkpoint behind each score in its final table, and the MOPD cards give no scores, so the blog post’s claim that benchmark performance “held steady” after the repair cannot be checked from the published numbers.

The card for the 9B model claims the fine-tune lifts Qwen3.5-9B from 32.0 to 44.6 on SWE-bench Pro and from 60.0 to 61.1 on SWE-bench Verified. The technical report lists a further version, trained with reinforcement learning, at 47.6 and 66.2. That version is not in the organisation’s repository list. “We release this SFT checkpoint as a starting point for open research in agentic reinforcement learning,” the card says. These are Xiaomi’s figures, and the 24.09 tokens a second above says nothing about them.

The cards state no hardware requirement and do not say whether a build for llama.cpp or MLX is planned. Xiaomi’s launch page returned no article text to an automated request, and we have not relied on it.

For scale, Aleph Alpha’s Kolibri, another open-weight release we have reported on, comes to 78.84 GB of weight files.

Sources

  1. Xiaomi, MiMo-V2.6-Pro-RL model card and technical report
  2. Xiaomi, MiMo-V2.6-Distill-Qwen-9B model card
Cite this page

Hill. "MiMo V2.6 is open weights from Xiaomi, and its 9B model runs at 24.09 tokens a second on a 48 GB Mac". Onticpost, published 5 October 2026. https://onticpost.com/xiaomi-mimo-v2-6-open-weights/

Updated 5 October 2026. This address stays the same when the piece is updated.

Share this