Model reference · open weights
Qwen3.5-Reverse-Text-RL is an open-weight language model from PrimeIntellect. Qwen3.5-0.8B-Reverse-Text-RL (BF16) weighs 2.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the PrimeIntellect/Qwen3.5-0.8B-Reverse-Text-RL model card, 2026-10-05
What it is
| Released by | PrimeIntellect |
|---|---|
| Released | 2026-10-04 |
| Parameters | 1.1B |
| VRAM | 2.2 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB | 52 | 16 | all 256K | 11.6 GB |
| RTX 4060 Ti 16 GB | 79 | 25 | all 256K | 15.4 GB |
| RTX 3090 24 GB | 137 | 43 | all 256K | 23.4 GB |
| RTX 4090 24 GB | 136 | 43 | all 256K | 23.4 GB |
| RTX 5090 32 GB | 191 | 60 | all 256K | 31.0 GB |
| L40S 48 GB | 284 | 89 | all 256K | 44.0 GB |
| A100 80 GB | 529 | 167 | all 256K | 78.2 GB |
| H100 80 GB | 492 | 155 | all 256K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 604 | 191 | all 256K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 701 | 221 | all 256K | 107 GB |
| H200 141 GB | 920 | 291 | all 256K | 138 GB |
| B200 180 GB | 1000+ | 377 | all 256K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 4.4 GB | 4.7 GB |
| 5 | 4.9 GB | 6.4 GB |
| 8 | 5.4 GB | 7.8 GB |
| 16 | 6.5 GB | 11.3 GB |
| 32 | 8.7 GB | 18.4 GB |
| 64 | 13.2 GB | 32.5 GB |
One card, with vLLM's small-card settings.
From the model card
A short RL fine-tune of PrimeIntellect/Qwen3.5-0.8B-Reverse-Text-SFT on the reverse-text environment with prime-rl.
It is meant as the frozen teacher / sampler for the prime-rl CI tests of on-policy distillation and RL-SFT once they move to Qwen3.5 (today they use PrimeIntellect/Qwen3-0.6B-Reverse-Text-RL). It is not meant for general use.
21814b401.examples/basic/reverse-text/rl.toml at that commit with model.name = the SFT checkpoint and orchestrator.renderer.name = "qwen3.5": 20 steps, 8 prompts x 16 rollouts (batch 128), 128 max tokens, lr 3e-6, 1 inference + 1 trainer H200, NCCL weight broadcast.reverse-text eval reward (256 prompts, temperature 1): step 0: 0.3086, step 5: 0.3543, step 10: 0.3967, step 15: 0.4027, step 20: 0.4077.On Qwen3.5-0.8B, prime-rl RL currently shows a much higher trainer/inference mismatch KL (0.01-0.1 per step) than on Qwen3-0.6B (about 0.002) with the same recipe, and learning stalls after about 20 steps. This is under investigation.
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.