Model reference · open weights

Muse-Glimmer-W4A4

LLMs Inferact Vision + text 1 build Open weights 41k dl/mo

Muse-Glimmer-W4A4 is an open-weight language model from Inferact. Muse-Glimmer-30B-NVFP4-W4A4 (NVFP4) weighs 25.4 GB; the smallest configuration that runs it is RTX 5090 32 GB.

What it is

Released byInferact
Released2026-08-10
Parameters6.0B
VRAM25.4 GB for the weights

What it runs on

Memory and cards for Muse-Glimmer-30B-NVFP4-W4A4 (NVFP4)

25.4 GBweights, file size
762 MBruntime overhead, at least

How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.

CardWeights alone
RTX 3060 12 GB … RTX 4090 24 GB
4 smaller cards
does not fit
RTX 5090 32 GBfits
L40S 48 GB
FP4 without its speed-up here
fits
A100 80 GB
FP4 without its speed-up here
fits
H100 80 GB
FP4 without its speed-up here
fits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GB
FP4 without its speed-up here
fits
B200 180 GBfits

From the model card

What Inferact says about Muse-Glimmer-W4A4

NVFP4 quantized weight for https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/README.md

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms