Model reference · open weights

Qwen3.5-Reverse-Text-SFT

NEW · this week LLMs PrimeIntellect Vision + text 1 build Open weights 0 dl/mo

Qwen3.5-Reverse-Text-SFT is an open-weight language model from PrimeIntellect. Qwen3.5-0.8B-Reverse-Text-SFT (BF16) weighs 2.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.

  • Qwen3.5-Reverse-Text-SFT is a 1.1B parameter image-text-to-text model by PrimeIntellect, created as a short supervised fine-tune of Qwen3.5-0.8B on the Reverse-Text-SFT dataset.
  • It serves as a deliberately under-trained starting point for reverse-text reinforcement learning tests and is not intended for general use.
  • The model supports a context length of 262144 tokens and is released under the apache-2.0 licence.

Summary of the PrimeIntellect/Qwen3.5-0.8B-Reverse-Text-SFT model card, 2026-10-05

What it is

Released byPrimeIntellect
Released2026-10-04
Parameters1.1B
VRAM2.2 GB for the weights

What it runs on

Memory and cards for Qwen3.5-0.8B-Reverse-Text-SFT (BF16)

2.2 GBweights, file size
12 MBcache per 1K tokens
39 MBfixed state per request
2.0 GBruntime overhead, at least
262,144 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB5216all 256K11.6 GB
RTX 4060 Ti 16 GB7925all 256K15.4 GB
RTX 3090 24 GB13743all 256K23.4 GB
RTX 4090 24 GB13643all 256K23.4 GB
RTX 5090 32 GB19160all 256K31.0 GB
L40S 48 GB28489all 256K44.0 GB
A100 80 GB529167all 256K78.2 GB
H100 80 GB492155all 256K78.1 GB
RTX PRO 6000 Blackwell 96 GB604191all 256K93.8 GB
DGX Spark (GB10) 128 GB unified701221all 256K107 GB
H200 141 GB920291all 256K138 GB
B200 180 GB1000+377all 256K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
14.4 GB4.7 GB
54.9 GB6.4 GB
85.4 GB7.8 GB
166.5 GB11.3 GB
328.7 GB18.4 GB
6413.2 GB32.5 GB

One card, with vLLM's small-card settings.

From the model card

What PrimeIntellect says about Qwen3.5-Reverse-Text-SFT

Read the model card

A short SFT fine-tune of Qwen/Qwen3.5-0.8B on PrimeIntellect/Reverse-Text-SFT. It is a small, deliberately under-trained starting point for reverse-text RL, made for the planned move of the prime-rl CI RL tests from PrimeIntellect/Qwen3-0.6B-Reverse-Text-SFT to Qwen3.5. It is not meant for general use.

Recipe

  • Code: prime-rl commit 21814b401 (branch ci/qwen3_5-ci, contains the Qwen3.5 tied lm_head fix #3863 and the CP fix #3864).
  • Config (uv run sft @ sft.toml), 1 H200:
max_steps = 10

[model]
name = "Qwen/Qwen3.5-0.8B"

[data]
name = "PrimeIntellect/Reverse-Text-SFT"
seq_len = 4096
batch_size = 32

[optim]
lr = 2e-5
  • Chat template: unchanged Qwen3.5 template, thinking off (the 0.8B default). Completions are rendered with the empty \n\n\n\n prefix, the same as RL generation.
  • Weights include the (frozen, unchanged) vision tower, so the checkpoint loads as Qwen3_5ForConditionalGeneration like the base model.

Numbers

  • SFT loss, steps 1-10: 4.98, 6.16, 5.59, 4.96, 4.53, 4.22, 3.96, 3.62, 3.27, 2.92.
  • reverse-text eval reward (LCS ratio, 256 prompts, temperature 1, 128 max tokens): 0.31 (base model: 0.03 on 32 prompts).

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms