Model reference · open weights
MiMo-Pro-MOPD is an open-weight language model from XiaomiMiMo. MiMo-V2.6-Pro-MOPD (FP8) weighs 566 GB; the smallest configuration that runs it is 8× A100 80 GB.
What it is
| Released by | XiaomiMiMo |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 1024.2B |
| Context | 1,048,576 tokens |
| Runs with | transformers |
| Released | 2026-09-27 |
| Popularity | 763 downloads / month |
| Weights | 566 GB (MiMo-V2.6-Pro-MOPD (FP8), file size) |
| Licence | Open weights |
What it runs on
Weights 566 GB (file size) · KV cache 430 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 3.5 GB on a small card · context up to 1,048,576 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … 8× H100 80 GB 13 smaller cards | — | — | — | |
| 8× A100 80 GB tensor parallel · FP8 without its speed-up here | 8 | 2 | 71K | 78.2 GB a card |
| 4× B200 180 GB tensor parallel | 18 | 4 | 146K | 176 GB a card |
| 8× H200 141 GB tensor parallel | 125 | 31 | 1000K | 138 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 573 GB | 584 GB |
| 5 | 587 GB | 640 GB |
| 8 | 598 GB | 682 GB |
| 16 | 626 GB | 795 GB |
| 32 | 682 GB | 1021 GB |
| 64 | 795 GB | 1471 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
| | | | | |
| | |
[!IMPORTANT] This is the MOPD upgrade of the MiMo-V2.6-Pro-RL checkpoint.
MOPD2 distills several domain-specialized teachers into the student on-policy. The teachers fall into two families: mixRL teachers, trained on verifiable tasks, and SFT teachers, trained on synthetic demonstrations for open-domain tasks where a reliable reward is hard to design. Three streams contribute to a single update:
Method details are in Technical Report §5.6.
Following the release of MiMo-V2.6, tool-call repetition emerged as one of the most noticeable issues in agentic settings: the model would sometimes issue the same or highly similar tool calls repeatedly, consuming time and context without making progress. This checkpoint mitigates it.
Figure: response-level repetition rate on MiMo-V2.6-Pro, RL-stage versus this checkpoint, across context lengths and agent harnesses.
The technical blog has the full diagnosis. The fix is lightweight to train: a short specialized-teacher run that folds into the normal MOPD pass.
Figure 1. MiMo-V2.6 architecture.
| Model | Download |
|---|---|
| MiMo-V2.6-Pro-RL | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Flash-RL | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Pro-MOPD | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Flash-MOPD | 🤗 HuggingFace · 🤖 ModelScope |
| Component | MiMo-V2.6-Pro-MOPD |
|---|---|
| Layers (Total / SWA / GA) | 70 / 60 / 10 |
| Hidden Size | 6144 |
| SWA Heads (Q/KV) | 128 / 8 |
| GA Heads (Q/KV) | 128 / 8 |
| Head Dimensions (QK / V) | 192 / 128 |
| Sliding Window Size | 128 |
| Routed Experts (Total / Activated) | 384 / 8 |
| Max Context Length | 1M |
| MTP / Speculative Decoder | 5 SWA layers, window 1024 |
The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.
| Configuration | Value |
|---|---|
| Layers (Total / SWA / GA) | 28 / 24 / 4 |
| Hidden Size | 1280 |
| Attention Heads (Q / KV) | 32 / 8 |
| Head Dimension | 64 |
| Patch Size (T × H × W) | 2 × 16 × 16 |
| Sliding Window (Left / Right) | 64 / 64 |
| Spatial Merge Size | 2 × 2 |
| Parameters | 681M |
AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).
5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.
For best performance, follow the SGLang MiMo cookbook. Docker image: lmsysorg/sglang:latest.
sglang serve \
--trust-remote-code \
--model-path XiaomiMiMo/MiMo-V2.6-Pro-MOPD \
--tp 16 \
--dp 2 \
--enable-dp-attention \
--mm-enable-dp-encoder \
--ep 16 \
--moe-a2a-backend deepep \
--moe-dense-tp-size 1 \
--mem-fraction-static 0.7 \
--max-running-requests 128 \
--chunked-prefill-size 32768 \
--page-size 64 \
--swa-full-tokens-ratio 0.3 \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--enable-multi-layer-eagle \
--reasoning-parser mimo \
--tool-call-parser mimo \
--hosQuoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.