Model reference · open weights
MiMo-Flash-MOPD is an open-weight language model from XiaomiMiMo. MiMo-V2.6-Flash-MOPD (FP8) weighs 173 GB; the smallest configuration that runs it is 2× H200 141 GB.
What it is
| Released by | XiaomiMiMo |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 310.8B |
| Context | 1,048,576 tokens |
| Runs with | transformers |
| Released | 2026-09-27 |
| Popularity | 637 downloads / month |
| Weights | 173 GB (MiMo-V2.6-Flash-MOPD (FP8), file size) |
| Licence | Open weights |
What it runs on
Weights 173 GB (file size) · KV cache 147 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 3.1 GB on a small card · context up to 1,048,576 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … B200 180 GB 12 smaller cards | — | — | — | |
| 2× H200 141 GB tensor parallel | 67 | 16 | 542K | 138 GB a card |
| 4× H100 80 GB tensor parallel | 80 | 20 | 644K | 78.1 GB a card |
| 4× A100 80 GB tensor parallel · FP8 without its speed-up here | 105 | 26 | 844K | 78.2 GB a card |
| 2× B200 180 GB tensor parallel | 123 | 30 | 984K | 176 GB a card |
| 4× RTX PRO 6000 Blackwell 96 GB tensor parallel | 132 | 33 | all 1024K | 93.8 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 177 GB | 181 GB |
| 5 | 182 GB | 200 GB |
| 8 | 186 GB | 215 GB |
| 16 | 195 GB | 253 GB |
| 32 | 215 GB | 331 GB |
| 64 | 253 GB | 485 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
| | | | | |
| | |
[!IMPORTANT] This is the MOPD upgrade of the MiMo-V2.6-Flash-RL checkpoint.
MOPD2 distills several domain-specialized teachers into the student on-policy. The teachers fall into two families: mixRL teachers, trained on verifiable tasks, and SFT teachers, trained on synthetic demonstrations for open-domain tasks where a reliable reward is hard to design. Three streams contribute to a single update:
Method details are in Technical Report §5.6.
Following the release of MiMo-V2.6, tool-call repetition emerged as one of the most noticeable issues in agentic settings: the model would sometimes issue the same or highly similar tool calls repeatedly, consuming time and context without making progress. This checkpoint mitigates it.
Figure: response-level repetition rate on MiMo-V2.6-Flash, RL-stage versus this checkpoint, across context lengths and agent harnesses.
The technical blog has the full diagnosis. The fix is lightweight to train: a short specialized-teacher run that folds into the normal MOPD pass.
Figure 1. MiMo-V2.6 architecture.
| Model | Download |
|---|---|
| MiMo-V2.6-Pro-RL | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Flash-RL | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Pro-MOPD | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Flash-MOPD | 🤗 HuggingFace · 🤖 ModelScope |
| Component | MiMo-V2.6-Flash-MOPD |
|---|---|
| Layers (Total / SWA / GA) | 48 / 39 / 9 |
| Hidden Size | 4096 |
| SWA Heads (Q/KV) | 64 / 8 |
| GA Heads (Q/KV) | 64 / 4 |
| Head Dimensions (QK / V) | 192 / 128 |
| Sliding Window Size | 128 |
| Routed Experts (Total / Activated) | 256 / 8 |
| Max Context Length | 1M |
| MTP / Speculative Decoder | 5 SWA layers, window 1024 |
The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.
| Configuration | Value |
|---|---|
| Layers (Total / SWA / GA) | 28 / 24 / 4 |
| Hidden Size | 1280 |
| Attention Heads (Q / KV) | 32 / 8 |
| Head Dimension | 64 |
| Patch Size (T × H × W) | 2 × 16 × 16 |
| Sliding Window (Left / Right) | 64 / 64 |
| Spatial Merge Size | 2 × 2 |
| Parameters | 681M |
AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).
5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.
For best performance, follow the SGLang MiMo cookbook. Docker image: lmsysorg/sglang:latest.
sglang serve \
--trust-remote-code \
--model-path XiaomiMiMo/MiMo-V2.6-Flash-MOPD \
--tp 8 \
--dp 2 \
--enable-dp-attention \
--enable-dp-lm-head \
--mm-enable-dp-encoder \
--mem-fraction-static 0.65 \
--chunked-prefill-size 16384 \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--enable-multi-layer-eagle \
--reasoning-parser mimo \
--tool-call-parser mimo \
--host 0.0.0.0 \
--port 30000
Follow the [vLLM MiMo-V2.5 recipe](https://recipes.vllm.ai/XiaomiMiMo/MiMo-
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.