Model reference · open weights
MiniCPM-MoE-8x2B is an open-weight language model from openbmb. MiniCPM-MoE-8x2B (BF16) weighs 27.7 GB; the smallest configuration that runs it is 2× RTX 4060 Ti 16 GB.
MiniCPM-MoE-8x2B is a decoder-only transformer-based generative language model developed by openbmb. It utilizes a Mixture-of-Experts architecture with 8 experts per layer, activating 2 for each token, and supports a context length of 4096 tokens. The model is instruction-tuned and available in bfloat16 precision.
Summary of the openbmb/MiniCPM-MoE-8x2B model card, 2026-10-01
What it is
| Released by | openbmb |
|---|---|
| Type | Language models |
| Task | Text gen · MoE |
| Context | 4,096 tokens |
| Runs with | transformers |
| Released | 2024-04-07 |
| Popularity | 4k downloads / month |
| Weights | 27.7 GB (MiniCPM-MoE-8x2B (BF16), file size) |
| Licence | Licence not stated |
What it runs on
Weights 27.7 GB (file size) · KV cache 369 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 640 MB on a small card · context up to 4,096 tokens.
| Card | Requests at once 4K, its whole window tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … RTX 4090 24 GB 4 smaller cards | — | — | — | |
| RTX 5090 32 GB | 1 | — | all 4K | 31.0 GB |
| L40S 48 GB | 10 | — | all 4K | 44.0 GB |
| A100 80 GB | 32 | — | all 4K | 78.2 GB |
| H100 80 GB | 30 | — | all 4K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 40 | — | all 4K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 49 | — | all 4K | 107 GB |
| H200 141 GB | 70 | — | all 4K | 138 GB |
| B200 180 GB | 95 | — | all 4K | 176 GB |
| 2× RTX 4060 Ti 16 GB tensor parallel | 1 | — | all 4K | 15.4 GB a card |
| 2× RTX 4090 24 GB tensor parallel | 11 | — | all 4K | 23.4 GB a card |
| 2× RTX 3090 24 GB tensor parallel | 11 | — | all 4K | 23.4 GB a card |
| 2× RTX 5090 32 GB tensor parallel | 21 | — | all 4K | 31.0 GB a card |
| Requests at once | 4K, its whole window tokens each | 32K tokens each |
|---|---|---|
| 1 | 29.9 GB | — |
| 5 | 35.9 GB | — |
| 8 | 40.5 GB | — |
| 16 | 52.5 GB | — |
| 32 | 76.7 GB | — |
| 64 | 125 GB | — |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
The MiniCPM-MoE-8x2B is a decoder-only transformer-based generative language model.
The MiniCPM-MoE-8x2B adopt a Mixture-of-Experts(MoE) architecture, which has 8 experts per layer and activates 2 of 8 experts for each token.
This is a model version after instruction tuning but without other rlhf methods. Chat template is automatically applied.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
torch.manual_seed(0)
path = 'openbmb/MiniCPM-MoE-8x2B'
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForCausalLM.from_pretrained(path, torch_dtype=torch.bfloat16, device_map='cuda', trust_remote_code=True)
responds, history = model.chat(tokenizer, "山东省最高的山是哪座山, 它比黄山高还是矮?差距多少?", temperature=0.8, top_p=0.8)
print(responds)
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.