Model reference · open weights
Muse-Glimmer-W4A4 is an open-weight language model from Inferact. Muse-Glimmer-30B-NVFP4-W4A4 (NVFP4) weighs 25.4 GB; the smallest configuration that runs it is RTX 5090 32 GB.
What it is
| Released by | Inferact |
|---|---|
| Released | 2026-08-10 |
| Parameters | 6.0B |
| VRAM | 25.4 GB for the weights |
What it runs on
How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.
| Card | Weights alone |
|---|---|
| RTX 3060 12 GB … RTX 4090 24 GB 4 smaller cards | does not fit |
| RTX 5090 32 GB | fits |
| L40S 48 GB FP4 without its speed-up here | fits |
| A100 80 GB FP4 without its speed-up here | fits |
| H100 80 GB FP4 without its speed-up here | fits |
| RTX PRO 6000 Blackwell 96 GB | fits |
| DGX Spark (GB10) 128 GB unified | fits |
| H200 141 GB FP4 without its speed-up here | fits |
| B200 180 GB | fits |
From the model card
NVFP4 quantized weight for https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/README.md
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.
How it works