Model reference · open weights

Qwen3-VL

LLMs amd Vision + text · MoE 1 build Open weights 23k dl/mo

Qwen3-VL is an open-weight language model from amd. Qwen3-VL-235B-A22B-Instruct-MXFP4 (MXFP4) weighs 135 GB in vLLM, 127 GB of files; the smallest configuration that runs it is B200 180 GB.

  • Qwen3-VL is a vision-language model developed by AMD for image-text-to-text tasks, featuring a 118.8B parameter MoE architecture.
  • It supports a native 256K context length expandable to 1M and handles OCR in 32 languages.
  • The model is released under the Apache-2.0 license.

Summary of the amd/Qwen3-VL-235B-A22B-Instruct-MXFP4 model card, 2026-10-01

What it is

Released byamd
Released2026-07-27
Parameters118.8B
VRAM135 GB in vLLM, 127 GB of files for the weights

What it runs on

Memory and cards for Qwen3-VL-235B-A22B-Instruct-MXFP4 (MXFP4)

135 GB in vLLM, 127 GB of filesweights, file size
762 MBruntime overhead, at least

How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.

CardWeights alone
RTX 3060 12 GB … H200 141 GB
11 smaller cards
does not fit
B200 180 GBfits
8× H100 80 GB
tensor parallel
does not fit
8× A100 80 GB
tensor parallel
does not fit
8× H200 141 GB
tensor parallel
does not fit

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms