Model reference · open weights

Qwen3.5-Reverse-Text-RL

NEW · this week LLMs PrimeIntellect Vision + text 1 build Open weights 0 dl/mo

Qwen3.5-Reverse-Text-RL is an open-weight language model from PrimeIntellect. Qwen3.5-0.8B-Reverse-Text-RL (BF16) weighs 2.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.

  • Qwen3.5-Reverse-Text-RL is a 1.1B parameter model by PrimeIntellect designed for image-text-to-text tasks.
  • It serves as a frozen teacher for continuous integration tests in the prime-rl framework and is not intended for general use.
  • The model supports a context length of 262144 tokens and is released under the apache-2.0 licence.

Summary of the PrimeIntellect/Qwen3.5-0.8B-Reverse-Text-RL model card, 2026-10-05

What it is

Released byPrimeIntellect
Released2026-10-04
Parameters1.1B
VRAM2.2 GB for the weights

What it runs on

Memory and cards for Qwen3.5-0.8B-Reverse-Text-RL (BF16)

2.2 GBweights, file size
12 MBcache per 1K tokens
39 MBfixed state per request
2.0 GBruntime overhead, at least
262,144 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB5216all 256K11.6 GB
RTX 4060 Ti 16 GB7925all 256K15.4 GB
RTX 3090 24 GB13743all 256K23.4 GB
RTX 4090 24 GB13643all 256K23.4 GB
RTX 5090 32 GB19160all 256K31.0 GB
L40S 48 GB28489all 256K44.0 GB
A100 80 GB529167all 256K78.2 GB
H100 80 GB492155all 256K78.1 GB
RTX PRO 6000 Blackwell 96 GB604191all 256K93.8 GB
DGX Spark (GB10) 128 GB unified701221all 256K107 GB
H200 141 GB920291all 256K138 GB
B200 180 GB1000+377all 256K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
14.4 GB4.7 GB
54.9 GB6.4 GB
85.4 GB7.8 GB
166.5 GB11.3 GB
328.7 GB18.4 GB
6413.2 GB32.5 GB

One card, with vLLM's small-card settings.

From the model card

What PrimeIntellect says about Qwen3.5-Reverse-Text-RL

Read the model card

A short RL fine-tune of PrimeIntellect/Qwen3.5-0.8B-Reverse-Text-SFT on the reverse-text environment with prime-rl. It is meant as the frozen teacher / sampler for the prime-rl CI tests of on-policy distillation and RL-SFT once they move to Qwen3.5 (today they use PrimeIntellect/Qwen3-0.6B-Reverse-Text-RL). It is not meant for general use.

Recipe

  • Code: prime-rl commit 21814b401.
  • Config: examples/basic/reverse-text/rl.toml at that commit with model.name = the SFT checkpoint and orchestrator.renderer.name = "qwen3.5": 20 steps, 8 prompts x 16 rollouts (batch 128), 128 max tokens, lr 3e-6, 1 inference + 1 trainer H200, NCCL weight broadcast.

Numbers

  • Train reward per step: 0.30, 0.30, 0.30, 0.32, 0.36, 0.33, 0.35, 0.34, 0.35, 0.35, 0.37, 0.40, 0.37, 0.36, 0.39, 0.38, 0.41, 0.43, 0.41, 0.42.
  • reverse-text eval reward (256 prompts, temperature 1): step 0: 0.3086, step 5: 0.3543, step 10: 0.3967, step 15: 0.4027, step 20: 0.4077.

Caveat

On Qwen3.5-0.8B, prime-rl RL currently shows a much higher trainer/inference mismatch KL (0.01-0.1 per step) than on Qwen3-0.6B (about 0.002) with the same recipe, and learning stalls after about 20 steps. This is under investigation.

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms