Model reference · open weights
Qwen3.5-Reverse-Text-SFT is an open-weight language model from PrimeIntellect. Qwen3.5-0.8B-Reverse-Text-SFT (BF16) weighs 2.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the PrimeIntellect/Qwen3.5-0.8B-Reverse-Text-SFT model card, 2026-10-05
What it is
| Released by | PrimeIntellect |
|---|---|
| Released | 2026-10-04 |
| Parameters | 1.1B |
| VRAM | 2.2 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB | 52 | 16 | all 256K | 11.6 GB |
| RTX 4060 Ti 16 GB | 79 | 25 | all 256K | 15.4 GB |
| RTX 3090 24 GB | 137 | 43 | all 256K | 23.4 GB |
| RTX 4090 24 GB | 136 | 43 | all 256K | 23.4 GB |
| RTX 5090 32 GB | 191 | 60 | all 256K | 31.0 GB |
| L40S 48 GB | 284 | 89 | all 256K | 44.0 GB |
| A100 80 GB | 529 | 167 | all 256K | 78.2 GB |
| H100 80 GB | 492 | 155 | all 256K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 604 | 191 | all 256K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 701 | 221 | all 256K | 107 GB |
| H200 141 GB | 920 | 291 | all 256K | 138 GB |
| B200 180 GB | 1000+ | 377 | all 256K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 4.4 GB | 4.7 GB |
| 5 | 4.9 GB | 6.4 GB |
| 8 | 5.4 GB | 7.8 GB |
| 16 | 6.5 GB | 11.3 GB |
| 32 | 8.7 GB | 18.4 GB |
| 64 | 13.2 GB | 32.5 GB |
One card, with vLLM's small-card settings.
From the model card
A short SFT fine-tune of Qwen/Qwen3.5-0.8B on PrimeIntellect/Reverse-Text-SFT.
It is a small, deliberately under-trained starting point for reverse-text RL, made for the planned move of the prime-rl CI RL tests from PrimeIntellect/Qwen3-0.6B-Reverse-Text-SFT to Qwen3.5. It is not meant for general use.
21814b401 (branch ci/qwen3_5-ci, contains the Qwen3.5 tied lm_head fix #3863 and the CP fix #3864).uv run sft @ sft.toml), 1 H200:max_steps = 10
[model]
name = "Qwen/Qwen3.5-0.8B"
[data]
name = "PrimeIntellect/Reverse-Text-SFT"
seq_len = 4096
batch_size = 32
[optim]
lr = 2e-5
\n\n\n\n prefix, the same as RL generation.Qwen3_5ForConditionalGeneration like the base model.reverse-text eval reward (LCS ratio, 256 prompts, temperature 1, 128 max tokens): 0.31 (base model: 0.03 on 32 prompts).Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.