Model reference · open weights
Qwen1.5-MoE is an open-weight language model from Qwen. Qwen1.5-MoE-A2.7B (BF16) weighs 28.6 GB; the smallest configuration that runs it is L40S 48 GB.
Qwen1.5-MoE is a transformer-based Mixture of Experts decoder-only language model developed by Qwen for text generation. The model contains 14.3B total parameters with 2.7B activated parameters during runtime and supports a context length of 8192 tokens. It is trained on English data and released under an other licence.
Summary of the Qwen/Qwen1.5-MoE-A2.7B model card, 2026-10-01
What it is
| Released by | Qwen |
|---|---|
| Type | Language models |
| Task | Text gen · MoE |
| Parameters (lead) | 14.3B |
| Context | 8,192 tokens |
| Runs with | transformers |
| Released | 2024-02-29 |
| Popularity | 716k downloads / month |
| Weights | 28.6 GB (Qwen1.5-MoE-A2.7B (BF16), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 28.6 GB (file size) · KV cache 197 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 1.0 GB on a small card · context up to 8,192 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … RTX 4090 24 GB 4 smaller cards | — | — | — | |
| RTX 5090 32 GB | — | — | 6K | 31.0 GB |
| L40S 48 GB | 8 | — | all 8K | 44.0 GB |
| A100 80 GB | 30 | — | all 8K | 78.2 GB |
| H100 80 GB | 26 | — | all 8K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 36 | — | all 8K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 45 | — | all 8K | 107 GB |
| H200 141 GB | 64 | — | all 8K | 138 GB |
| B200 180 GB | 86 | — | all 8K | 176 GB |
| 2× RTX 4090 24 GB tensor parallel | 9 | — | all 8K | 23.4 GB a card |
| 2× RTX 3090 24 GB tensor parallel | 10 | — | all 8K | 23.4 GB a card |
| 2× RTX 5090 32 GB tensor parallel | 19 | — | all 8K | 31.0 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 31.2 GB | — |
| 5 | 37.7 GB | — |
| 8 | 42.5 GB | — |
| 16 | 55.4 GB | — |
| 32 | 81.2 GB | — |
| 64 | 133 GB | — |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
Qwen1.5-MoE is a transformer-based MoE decoder-only language model pretrained on a large amount of data.
For more details, please refer to our blog post and GitHub repo.
Qwen1.5-MoE employs Mixture of Experts (MoE) architecture, where the models are upcycled from dense language models. For instance, Qwen1.5-MoE-A2.7B is upcycled from Qwen-1.8B. It has 14.3B parameters in total and 2.7B activated parameters during runtime, while achieving comparable performance to Qwen1.5-7B, it only requires 25% of the training resources. We also observed that the inference speed is 1.74 times that of Qwen1.5-7B.
The code of Qwen1.5-MoE has been in the latest Hugging face transformers and we advise you to build from source with command pip install git+https://github.com/huggingface/transformers, or you might encounter the following error:
KeyError: 'qwen2_moe'.
We do not advise you to use base language models for text generation. Instead, you can apply post-training, e.g., SFT, RLHF, continued pretraining, etc., on this model.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.