Model reference · open weights
Qwen3-untied is an open-weight language model from PrimeIntellect. Qwen3-0.6B-untied (BF16) weighs 1.5 GB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the PrimeIntellect/Qwen3-0.6B-untied model card, 2026-10-05
What it is
| Released by | PrimeIntellect |
|---|---|
| Released | 2026-10-05 |
| Parameters | 752M |
| VRAM | 1.5 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB | 10 | 2 | all 40K | 11.6 GB |
| RTX 4060 Ti 16 GB | 14 | 3 | all 40K | 15.4 GB |
| RTX 3090 24 GB | 22 | 5 | all 40K | 23.4 GB |
| RTX 4090 24 GB | 22 | 5 | all 40K | 23.4 GB |
| RTX 5090 32 GB | 30 | 7 | all 40K | 31.0 GB |
| L40S 48 GB | 44 | 11 | all 40K | 44.0 GB |
| A100 80 GB | 80 | 20 | all 40K | 78.2 GB |
| H100 80 GB | 76 | 19 | all 40K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 93 | 23 | all 40K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 107 | 26 | all 40K | 107 GB |
| H200 141 GB | 140 | 35 | all 40K | 138 GB |
| B200 180 GB | 180 | 45 | all 40K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 3.1 GB | 5.9 GB |
| 5 | 6.8 GB | 20.9 GB |
| 8 | 9.6 GB | 32.2 GB |
| 16 | 17.2 GB | 62.3 GB |
| 32 | 32.2 GB | 122 GB |
| 64 | 62.3 GB | 243 GB |
One card, with vLLM's small-card settings.
From the model card
Untied copy of PrimeIntellect/Qwen3-0.6B. Same weights;
lm_head.weightis stored as a copy of the input embedding andtie_word_embeddingsisfalse, so trainers that don't support tied LM heads (e.g. prime-rl) can load it. Made withtools/untie_word_embeddings.pyfrom prime-rl.
This is a clone of Qwen/Qwen3-0.6B with a multi-turn, tool-call compatible chat template.
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.