Model reference · open weights
tinctura is an open-weight language model from bench-labs. tinctura-v1 (FP32) weighs 1.3 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | bench-labs |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 96M |
| Context | 2,048 tokens |
| Runs with | transformers |
| Released | 2026-09-23 |
| Popularity | 509 downloads / month |
| Weights | 1.3 GB (tinctura-v1 (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 1.3 GB (file size) · KV cache 23 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 486 MB on a small card · context up to 2,048 tokens.
| Card | Requests at once 2K, its whole window tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB | 207 | — | all 2K | 11.6 GB |
| RTX 4060 Ti 16 GB | 287 | — | all 2K | 15.4 GB |
| RTX 3090 24 GB | 457 | — | all 2K | 23.4 GB |
| RTX 4090 24 GB | 456 | — | all 2K | 23.4 GB |
| RTX 5090 32 GB | 618 | — | all 2K | 31.0 GB |
| L40S 48 GB | 892 | — | all 2K | 44.0 GB |
| A100 80 GB | 1000+ | — | all 2K | 78.2 GB |
| H100 80 GB | 1000+ | — | all 2K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 1000+ | — | all 2K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 1000+ | — | all 2K | 107 GB |
| H200 141 GB | 1000+ | — | all 2K | 138 GB |
| B200 180 GB | 1000+ | — | all 2K | 176 GB |
| Requests at once | 2K, its whole window tokens each | 32K tokens each |
|---|---|---|
| 1 | 1.9 GB | — |
| 5 | 2.1 GB | — |
| 8 | 2.2 GB | — |
| 16 | 2.6 GB | — |
| 32 | 3.3 GB | — |
| 64 | 4.9 GB | — |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.
From the model card
A 96.2M parameter decoder-only language model pretrained from scratch on 75B tokens. It is the cagliostro-v3 recipe at two thirds of the size: the same architecture, data, schedule and token count, with 18 layers instead of 30.
On the Open SLM Leaderboard Index it scores 20.81. Among models under 100M parameters that places it second, 0.26 behind Rose-1.5-Medium at 21.07.
The training run is complete. 75.00B tokens, 762,939 steps, learning rate decayed to zero.
Zero-shot, measured on the public weights in this repository with exactly the commands under Reproducing the evaluation: the lm-evaluation-harness hf backend and the leaderboard's official ArithMark-3 script, both in float32.
| Benchmark | Metric | Score |
|---|---|---|
| HellaSwag | acc_norm | 37.96 |
| ARC-Easy | acc_norm | 47.98 |
| ARC-Challenge | acc_norm | 25.77 |
| PIQA | acc_norm | 65.61 |
| ArithMark-3 | acc_norm | 38.40 |
| Open SLM Index | 20.81 |
The Index is the leaderboard's own formula, (N(HellaSwag,25) + N(CombinedARC,25) + N(PIQA,50) + 0.65*N(ArithMark,25)) / 3.65 where N(v,c) = 100(v-c)/(100-c) and CombinedARC is the mean of ARC-Easy and ARC-Challenge. Everything further down that tracks training uses the training-time evaluation harness. That harness tokenizes each answer separately from its question and ran ArithMark-3 in bfloat16, so it reads a little differently: 20.77 for the final checkpoint, with individual benchmarks up to 0.7 points apart from the table above. The numbers in this table are the ones to compare against the leaderboard.
The top of the sub-100M field, using the leaderboard's published figures. tinctura-v1 is scored with the same formula but is not yet listed on the board.
| Model | Params | Index |
|---|---|---|
| Rose-1.5-Medium | 98.2M | 21.07 |
| tinctura-v1 | 96.2M | 20.81 |
| Surjo-100m | 97.7M | 18.86 |
| cRia-LM-75M | 75.7M | 18.06 |
| Rose-Medium | 97.8M | 17.73 |
| Surjo-50M | 53.8M | 16.40 |
Against the leader the result is split. tinctura-v1 is ahead on ARC-Easy and PIQA, level on HellaSwag, and behind on ARC-Challenge and ArithMark-3. ArithMark-3 is most of the gap: 2.3 points there cost about 0.55 Index at its 0.65 weight, more than the whole 0.26 margin. Rose-1.5-Medium was trained on roughly 100B tokens according to its card, against 75B here.
Because tinctura-v1 follows cagliostro-v3 step for step, the two runs can be compared at identical points in training. Several of these comparisons were run side by side on the same machine with the same evaluation code, using checkpoints from both repositories. Every number in this section comes from the training-time harness, for both models.
Through the stable phase the smaller model trailed by about 1.7 Index over the first 20B tokens and by about 2.9 after that. Single evaluations carry roughly one point of noise, so individual gaps ranged from 0.5 to 3.9. Just before the cooldown, at step 646,000, the paired evaluation read 19.34 against 21.96.
The cooldown is where the two models part. cagliostro-v3 gained 4.6 Index over the last 15% of training. tinctura-v1 gained 1.4. The shortfall came in two places. From step 646,000 to 706,000 cagliostro-v3 added 1.95 and tinctura-v1 only 0.21. From 706,000 to 736,000 the two moved together, 1.32 and 1.23. Over the final stretch to the end cagliostro-v3 added another 1.32 while tinctura-v1 held flat. On ARC-Challenge and ArithMark-3 the smaller model finished the cooldown slightly lower than it started it.
After the first few billion tokens the training loss gap to cagliostro-v3 settled between 0.07 and 0.10 and stayed there through the cooldown. The sharp drop at 63.75B is the data mixture changing, not the model improving.
The first attempt at the cooldown hit a bug in the data loader. When training switches to the cooldown mixture, each process restarts its data-loading threads. Threads blocked on a full queue could miss the stop signal and keep producing batches from the old mixture next to the new ones, so each process trained on a blend. A standalone test of the loader leaked old-mixture batches in 9 of 10 trials, 20 to 44% of batches after the switch. The loss told on it: at the switch the loss fell by 0.26 where cagliostro-v3's fell by 0.49.
The loader was fixed so that every restart gets its own queue and old threads exit on their own, and the fixed version leaked nothing in the same test. The cooldown was then rerun from the last checkpoint before it, step 644,000. The rerun's loss fell by 0.45 at the switch and tracked cagliostro-v3 as expected. The weights in this repository come from the rerun.
The fix did not change the benchmarks in a measurable way. At step 706,000 the first attempt scored 19.68 and the rerun 19.55, which is within evaluation noise. The data-mixing bug was real, but it was not what held the cooldown gain down.
| Field | Value |
|---|---|
| Parameters | 96,200,064 |
| Non-embedding parameters | 78.2% |
| Layers | 18 |
| Hidden size | 640 |
| Intermediate size | 1,536 |
| Attention heads | 10 |
| Key/value heads | 5 |
| Attention | Grouped query attention with cross-head subspace attenuation |
| Activation | SwiGLU |
| Normalization | RMSNorm, eps 1e-6 |
| Positional encoding | RoPE, theta 100,000 |
| Context length | 2,048 |
| Vocabulary | 32,768 BPE |
| Embeddings | Tied input and output |
| Logit cap | 15.0 |
| Weights | float32 safetensors |
The architecture is defined in this repository and is the same code as cagliostro-v3. trust_remote_code=True is required because CagliostroForCausalLM is not part of transformers.
The same mixtures and shards as cagliostro-v3. The first mixture covers the stable phase, and the second takes over when the cooldown begins at 85% of
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.