Model reference · open weights
Darwin-RSI is an open-weight language model from FINAL-Bench. Darwin-27B-RSI (BF16) weighs 53.8 GB; the smallest configuration that runs it is H100 80 GB.
What it is
| Released by | FINAL-Bench |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 26.9B |
| Context | 262,144 tokens |
| Runs with | transformers |
| Based on | FINAL-Bench/Darwin-27B-Opus |
| Released | 2026-09-27 |
| Popularity | 619 downloads / month |
| Weights | 53.8 GB (Darwin-27B-RSI (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 53.8 GB (file size) · KV cache 66 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · 308 MB of fixed state per request · runtime overhead from 910 MB on a small card · context up to 262,144 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … L40S 48 GB 6 smaller cards | — | — | — | |
| A100 80 GB | 27 | 9 | all 256K | 78.2 GB |
| H100 80 GB | 21 | 7 | all 256K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 39 | 13 | all 256K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 55 | 19 | all 256K | 107 GB |
| H200 141 GB | 91 | 31 | all 256K | 138 GB |
| B200 180 GB | 136 | 46 | all 256K | 176 GB |
| 2× L40S 48 GB tensor parallel | 38 | 13 | all 256K | 44.0 GB a card |
| 4× RTX 4090 24 GB tensor parallel | 42 | 14 | all 256K | 23.4 GB a card |
| 4× RTX 3090 24 GB tensor parallel | 42 | 14 | all 256K | 23.4 GB a card |
| 4× RTX 5090 32 GB tensor parallel | 78 | 27 | all 256K | 31.0 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 55.5 GB | 57.2 GB |
| 5 | 58.9 GB | 67.0 GB |
| 8 | 61.5 GB | 74.3 GB |
| 16 | 68.2 GB | 94.0 GB |
| 32 | 81.7 GB | 133 GB |
| 64 | 109 GB | 212 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (hybrid: linear attention with full attention every few layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
Qwen3.5-27B family · 27B dense · Thinking mode · BF16 · Apache 2.0 No human-written answers. The model generated its own learning signal — and got measurably better.
Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution).
Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests.
As the reasoning engine of Darwin-27B-JEV on the Decision Index, it lifts the hardest reasoning decisions: GPQA Diamond skill 0.31 → 0.71, GSM8K 0.61 → 0.97, MMLU-Pro 0.60 → 0.82.
Darwin-27B-RSI is Model-level RSI: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. You download a new model file, and it is smarter on its own.
Harness-level RSI (for example, Google's RRSI) improves the prompts, tools and workflow around a fixed model — like rewriting an employee's manual, while Model-level RSI is the employee getting smarter. The two are complementary: a harness-level loop can run on top of a Model-level RSI model.
Most models improve only when people write more answers for them. Recursive Self-Improvement removes that bottleneck: the model works on problems, judges its own work, and learns from what it produced — then repeats. Each improved model becomes the starting point for the next improvement.
Darwin-27B-RSI demonstrates this loop on a 27B model:
The training procedure itself is not released.
| Benchmark | Darwin-27B-Opus | Darwin-27B-RSI | Δ |
|---|---|---|---|
| GPQA Diamond (1 sample) | 72.85 | 78.09 | +5.24 |
| GPQA Diamond (majority@16) | 79.80 | 83.59 | +3.79 |
| SuperGPQA (1 sample) | +4.03 |
Both models were measured under the same protocol (single sample, identical sampling settings and token budget), so numbers differ from the Darwin-27B-Opus card, which reports a different protocol. All gains are statistically significant in paired tests.
The Decision Index scores typed-decision engines on 43 benchmarks and ~121K decisions (chance-corrected: 0 = random, 1 = perfect). Darwin-27B-RSI handles the decisions that need real thinking:
| Benchmark (skill) | before | with Darwin-27B-RSI |
|---|---|---|
| GPQA Diamond ★ | 0.31 | 0.71 |
| GSM8K | 0.61 | 0.97 |
| CRUXEval | 0.61 | 0.87 |
| MMLU-Pro ★ | 0.60 | 0.82 |
| BBH ★ | 0.68 | 0.83 |
| CLadder | 0.49 | 0.70 |
★ = gold benchmark (weighted 1.2× on the board). Darwin-27B-JEV: ≈ 61.1 under the v0.2.1 board rules (our recomputation; official score pending review). Full run: FINAL-Bench/Darwin-27B-JEV-decision-index.
Darwin-27B-RSI is a thinking model. Give it room to reason and read the answer after the reasoning block.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "FINAL-Bench/Darwin-27B-RSI"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto")
messages = [{"role": "user", "content": "A ball is thrown upward at 40 m/s. For how long is it above 40 m? (g = 10 m/s²) Think, then give the final answer."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=8192, temperature=0.6, top_p=0.95, do_sample=True)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
vllm serve FINAL-Bench/Darwin-27B-RSI --max-model-len 32768
Recommended: temperature 0.6, top_p 0.95, generous token budget (8K–16K) for hard problems.
| Parent | FINAL-Bench/Darwin-27B-Opus |
| Architecture | Qwen3.5 family, 27B dense |
| Precision | BF16 |
| Improvement method | Recursive Self-Improvement, no human labels |
| License | Apache 2.0 |
| Developer | VIDRAFT · FINAL-Bench |
max_new_tokens for latency-sensitive use.@misc{darwin27b_rsi_2026,
title = {Darwin-27B-RSI: Recursive Self-Improvement without Human Labels},
author = {VIDRAFT and FINAL-Bench},
year = {2026},
url = {https://huggingface.co/FINAL-Bench/Darwin-27B-RSI}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.