Model reference · open weights

MiMo-Pro-MOPD

NEW · this week LLMs XiaomiMiMo Text gen 1 build Open weights 763 dl/mo

MiMo-Pro-MOPD is an open-weight language model from XiaomiMiMo. MiMo-V2.6-Pro-MOPD (FP8) weighs 566 GB; the smallest configuration that runs it is 8× A100 80 GB.

What it is

Released byXiaomiMiMo
TypeLanguage models
TaskText gen
Parameters (lead)1024.2B
Context1,048,576 tokens
Runs withtransformers
Released2026-09-27
Popularity763 downloads / month
Weights566 GB (MiMo-V2.6-Pro-MOPD (FP8), file size)
LicenceOpen weights

What it runs on

Memory and cards for MiMo-V2.6-Pro-MOPD (FP8)

Weights 566 GB (file size) · KV cache 430 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 3.5 GB on a small card · context up to 1,048,576 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … 8× H100 80 GB
13 smaller cards
———
8× A100 80 GB
tensor parallel · FP8 without its speed-up here
8271K78.2 GB a card
4× B200 180 GB
tensor parallel
184146K176 GB a card
8× H200 141 GB
tensor parallel
125311000K138 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
1573 GB584 GB
5587 GB640 GB
8598 GB682 GB
16626 GB795 GB
32682 GB1021 GB
64795 GB1471 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What XiaomiMiMo says about MiMo-Pro-MOPD

|  |  |  |  |  |

 |   |   | 

MiMo-V2.6-Pro-MOPD

[!IMPORTANT] This is the MOPD upgrade of the MiMo-V2.6-Pro-RL checkpoint.

Read the full model card
  • MOPD2 (👉 Technical Report §5.6) Fuses several domain-specialized teachers into one model, extending to domains where reliable training-time verification is hard, such as long-horizon game development, scientific research and embodied intelligence.
  • Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6 (👉 Technical Blog) An easy-to-overlook failure mode in which the model keeps issuing the same or highly similar tool calls, appearing busy while making no progress. Nothing fails outright, so it tends to go unnoticed. The MOPD stage handles it efficiently, with a short specialized-teacher run that converges quickly.

1. Introduction

How MOPD2 works

MOPD2 distills several domain-specialized teachers into the student on-policy. The teachers fall into two families: mixRL teachers, trained on verifiable tasks, and SFT teachers, trained on synthetic demonstrations for open-domain tasks where a reliable reward is hard to design. Three streams contribute to a single update:

  • Standard MOPD: mixRL teachers supervise full autonomous rollouts.
  • Teacher-Prefix OPD: prefixes come from teacher rollouts. A trajectory with k assistant turns yields k history prefixes, one per turn. The model generates a single new turn from each, and the teacher scores it against the same history.
  • SFT-Prefix OPD: prefixes come from SFT demonstrations. The demonstration supplies the history, and the model writes its own continuation.

Method details are in Technical Report §5.6.

Tool-call repetition

Following the release of MiMo-V2.6, tool-call repetition emerged as one of the most noticeable issues in agentic settings: the model would sometimes issue the same or highly similar tool calls repeatedly, consuming time and context without making progress. This checkpoint mitigates it.

Figure: response-level repetition rate on MiMo-V2.6-Pro, RL-stage versus this checkpoint, across context lengths and agent harnesses.

The technical blog has the full diagnosis. The fix is lightweight to train: a short specialized-teacher run that folds into the normal MOPD pass.

Model Summary

  • Architecture: Sparse MoE (Mixture of Experts), 1.02T total / 42B activated parameters
  • Context Length: 1M tokens
  • Modalities: Text, Image, Video, Audio
  • Vision Encoder: 681M-param MiMo ViT (28 layers: 24 SWA + 4 Full)
  • Audio Encoder: 308M AudioTokenizer + 127M audio patch encoder
  • Multi-Token Prediction (MTP): 5-layer speculative decoder

Figure 1. MiMo-V2.6 architecture.

2. Downloads

ModelDownload
MiMo-V2.6-Pro-RL🤗 HuggingFace · 🤖 ModelScope
MiMo-V2.6-Flash-RL🤗 HuggingFace · 🤖 ModelScope
MiMo-V2.6-Pro-MOPD🤗 HuggingFace · 🤖 ModelScope
MiMo-V2.6-Flash-MOPD🤗 HuggingFace · 🤖 ModelScope

3. Model Architecture

LLM Backbone

ComponentMiMo-V2.6-Pro-MOPD
Layers (Total / SWA / GA)70 / 60 / 10
Hidden Size6144
SWA Heads (Q/KV)128 / 8
GA Heads (Q/KV)128 / 8
Head Dimensions (QK / V)192 / 128
Sliding Window Size128
Routed Experts (Total / Activated)384 / 8
Max Context Length1M
MTP / Speculative Decoder5 SWA layers, window 1024

The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.

Vision Encoder (MiMo ViT)

ConfigurationValue
Layers (Total / SWA / GA)28 / 24 / 4
Hidden Size1280
Attention Heads (Q / KV)32 / 8
Head Dimension64
Patch Size (T × H × W)2 × 16 × 16
Sliding Window (Left / Right)64 / 64
Spatial Merge Size2 × 2
Parameters681M

Audio Encoders

AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).

Speculative Decoder

5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.

4. Deployment

For best performance, follow the SGLang MiMo cookbook. Docker image: lmsysorg/sglang:latest.

SGLang

sglang serve \
  --trust-remote-code \
  --model-path XiaomiMiMo/MiMo-V2.6-Pro-MOPD \
  --tp 16 \
  --dp 2 \
  --enable-dp-attention \
  --mm-enable-dp-encoder \
  --ep 16 \
  --moe-a2a-backend deepep \
  --moe-dense-tp-size 1 \
  --mem-fraction-static 0.7 \
  --max-running-requests 128 \
  --chunked-prefill-size 32768 \
  --page-size 64 \
  --swa-full-tokens-ratio 0.3 \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --enable-multi-layer-eagle \
  --reasoning-parser mimo \
  --tool-call-parser mimo \
  --hos

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms