Model reference · open weights

Darwin-RSI

NEW · this week LLMs FINAL-Bench Text gen 1 build Open weights 619 dl/mo

Darwin-RSI is an open-weight language model from FINAL-Bench. Darwin-27B-RSI (BF16) weighs 53.8 GB; the smallest configuration that runs it is H100 80 GB.

What it is

Released byFINAL-Bench
TypeLanguage models
TaskText gen
Parameters (lead)26.9B
Context262,144 tokens
Runs withtransformers
Based onFINAL-Bench/Darwin-27B-Opus
Released2026-09-27
Popularity619 downloads / month
Weights53.8 GB (Darwin-27B-RSI (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for Darwin-27B-RSI (BF16)

Weights 53.8 GB (file size) · KV cache 66 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · 308 MB of fixed state per request · runtime overhead from 910 MB on a small card · context up to 262,144 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … L40S 48 GB
6 smaller cards
———
A100 80 GB279all 256K78.2 GB
H100 80 GB217all 256K78.1 GB
RTX PRO 6000 Blackwell 96 GB3913all 256K93.8 GB
DGX Spark (GB10) 128 GB unified5519all 256K107 GB
H200 141 GB9131all 256K138 GB
B200 180 GB13646all 256K176 GB
2× L40S 48 GB
tensor parallel
3813all 256K44.0 GB a card
4× RTX 4090 24 GB
tensor parallel
4214all 256K23.4 GB a card
4× RTX 3090 24 GB
tensor parallel
4214all 256K23.4 GB a card
4× RTX 5090 32 GB
tensor parallel
7827all 256K31.0 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
155.5 GB57.2 GB
558.9 GB67.0 GB
861.5 GB74.3 GB
1668.2 GB94.0 GB
3281.7 GB133 GB
64109 GB212 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (hybrid: linear attention with full attention every few layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What FINAL-Bench says about Darwin-RSI

Qwen3.5-27B family · 27B dense · Thinking mode · BF16 · Apache 2.0 No human-written answers. The model generated its own learning signal — and got measurably better.


Read the full model card

Abstract

Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution).

Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests.

As the reasoning engine of Darwin-27B-JEV on the Decision Index, it lifts the hardest reasoning decisions: GPQA Diamond skill 0.31 → 0.71, GSM8K 0.61 → 0.97, MMLU-Pro 0.60 → 0.82.


Model-level RSI vs. harness-level RSI

Darwin-27B-RSI is Model-level RSI: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. You download a new model file, and it is smarter on its own.

Harness-level RSI (for example, Google's RRSI) improves the prompts, tools and workflow around a fixed model — like rewriting an employee's manual, while Model-level RSI is the employee getting smarter. The two are complementary: a harness-level loop can run on top of a Model-level RSI model.

What Is RSI?

Most models improve only when people write more answers for them. Recursive Self-Improvement removes that bottleneck: the model works on problems, judges its own work, and learns from what it produced — then repeats. Each improved model becomes the starting point for the next improvement.

Darwin-27B-RSI demonstrates this loop on a 27B model:

  • Human-written solutions or reasoning traces used: 0
  • Direction of change: measurably better on held-out graduate-level science
  • Contamination check: training problems share 0 items with the evaluation sets reported here

The training procedure itself is not released.


Results

Science reasoning (same protocol for both models)

BenchmarkDarwin-27B-OpusDarwin-27B-RSIΔ
GPQA Diamond (1 sample)72.8578.09+5.24
GPQA Diamond (majority@16)79.8083.59+3.79
SuperGPQA (1 sample)+4.03

Both models were measured under the same protocol (single sample, identical sampling settings and token budget), so numbers differ from the Darwin-27B-Opus card, which reports a different protocol. All gains are statistically significant in paired tests.

Decision Index — as the reasoning engine of Darwin-27B-JEV

The Decision Index scores typed-decision engines on 43 benchmarks and ~121K decisions (chance-corrected: 0 = random, 1 = perfect). Darwin-27B-RSI handles the decisions that need real thinking:

Benchmark (skill)beforewith Darwin-27B-RSI
GPQA Diamond ★0.310.71
GSM8K0.610.97
CRUXEval0.610.87
MMLU-Pro ★0.600.82
BBH ★0.680.83
CLadder0.490.70

★ = gold benchmark (weighted 1.2× on the board). Darwin-27B-JEV: ≈ 61.1 under the v0.2.1 board rules (our recomputation; official score pending review). Full run: FINAL-Bench/Darwin-27B-JEV-decision-index.


Usage

Darwin-27B-RSI is a thinking model. Give it room to reason and read the answer after the reasoning block.

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "FINAL-Bench/Darwin-27B-RSI"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto")

messages = [{"role": "user", "content": "A ball is thrown upward at 40 m/s. For how long is it above 40 m? (g = 10 m/s²) Think, then give the final answer."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=8192, temperature=0.6, top_p=0.95, do_sample=True)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))

vLLM

vllm serve FINAL-Bench/Darwin-27B-RSI --max-model-len 32768

Recommended: temperature 0.6, top_p 0.95, generous token budget (8K–16K) for hard problems.


Model Details

ParentFINAL-Bench/Darwin-27B-Opus
ArchitectureQwen3.5 family, 27B dense
PrecisionBF16
Improvement methodRecursive Self-Improvement, no human labels
LicenseApache 2.0
DeveloperVIDRAFT · FINAL-Bench

Limitations and Disclosure

  • Gains were measured on graduate-level science; other domains may change less.
  • As a thinking model, it can produce long reasoning; cap max_new_tokens for latency-sensitive use.
  • 31 training problems (0.22% of the benchmark) overlap with the Decision Index MMLU set; the model learned only from its own solutions to them.
  • Not affiliated with TypeSafe AI or its Jev product.

Citation

@misc{darwin27b_rsi_2026,
  title  = {Darwin-27B-RSI: Recursive Self-Improvement without Human Labels},
  author = {VIDRAFT and FINAL-Bench},
  year   = {2026},
  url    = {https://huggingface.co/FINAL-Bench/Darwin-27B-RSI}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms