Model reference · open weights
Qwen3-VL is an open-weight language model from amd. Qwen3-VL-235B-A22B-Instruct-MXFP4 (MXFP4) weighs 135 GB in vLLM, 127 GB of files; the smallest configuration that runs it is B200 180 GB.
Summary of the amd/Qwen3-VL-235B-A22B-Instruct-MXFP4 model card, 2026-10-01
What it is
| Released by | amd |
|---|---|
| Released | 2026-07-27 |
| Parameters | 118.8B |
| VRAM | 135 GB in vLLM, 127 GB of files for the weights |
What it runs on
How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.
| Card | Weights alone |
|---|---|
| RTX 3060 12 GB … H200 141 GB 11 smaller cards | does not fit |
| B200 180 GB | fits |
| 8× H100 80 GB tensor parallel | does not fit |
| 8× A100 80 GB tensor parallel | does not fit |
| 8× H200 141 GB tensor parallel | does not fit |
How it works