Model reference · open weights

tinctura

LLMs bench-labs Text gen 1 build Open weights 509 dl/mo

tinctura is an open-weight language model from bench-labs. tinctura-v1 (FP32) weighs 1.3 GB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released bybench-labs
TypeLanguage models
TaskText gen
Parameters (lead)96M
Context2,048 tokens
Runs withtransformers
Released2026-09-23
Popularity509 downloads / month
Weights1.3 GB (tinctura-v1 (FP32), file size)
LicenceOpen weights

What it runs on

Memory and cards for tinctura-v1 (FP32)

Weights 1.3 GB (file size) · KV cache 23 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 486 MB on a small card · context up to 2,048 tokens.

CardRequests at once
2K, its whole window tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB207—all 2K11.6 GB
RTX 4060 Ti 16 GB287—all 2K15.4 GB
RTX 3090 24 GB457—all 2K23.4 GB
RTX 4090 24 GB456—all 2K23.4 GB
RTX 5090 32 GB618—all 2K31.0 GB
L40S 48 GB892—all 2K44.0 GB
A100 80 GB1000+—all 2K78.2 GB
H100 80 GB1000+—all 2K78.1 GB
RTX PRO 6000 Blackwell 96 GB1000+—all 2K93.8 GB
DGX Spark (GB10) 128 GB unified1000+—all 2K107 GB
H200 141 GB1000+—all 2K138 GB
B200 180 GB1000+—all 2K176 GB
Memory needed at each load
Requests at once2K, its whole window tokens each32K tokens each
11.9 GB—
52.1 GB—
82.2 GB—
162.6 GB—
323.3 GB—
644.9 GB—

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.

From the model card

What bench-labs says about tinctura

A 96.2M parameter decoder-only language model pretrained from scratch on 75B tokens. It is the cagliostro-v3 recipe at two thirds of the size: the same architecture, data, schedule and token count, with 18 layers instead of 30.

On the Open SLM Leaderboard Index it scores 20.81. Among models under 100M parameters that places it second, 0.26 behind Rose-1.5-Medium at 21.07.

The training run is complete. 75.00B tokens, 762,939 steps, learning rate decayed to zero.

Read the full model card

Results

Zero-shot, measured on the public weights in this repository with exactly the commands under Reproducing the evaluation: the lm-evaluation-harness hf backend and the leaderboard's official ArithMark-3 script, both in float32.

BenchmarkMetricScore
HellaSwagacc_norm37.96
ARC-Easyacc_norm47.98
ARC-Challengeacc_norm25.77
PIQAacc_norm65.61
ArithMark-3acc_norm38.40
Open SLM Index20.81

The Index is the leaderboard's own formula, (N(HellaSwag,25) + N(CombinedARC,25) + N(PIQA,50) + 0.65*N(ArithMark,25)) / 3.65 where N(v,c) = 100(v-c)/(100-c) and CombinedARC is the mean of ARC-Easy and ARC-Challenge. Everything further down that tracks training uses the training-time evaluation harness. That harness tokenizes each answer separately from its question and ran ArithMark-3 in bfloat16, so it reads a little differently: 20.77 for the final checkpoint, with individual benchmarks up to 0.7 points apart from the table above. The numbers in this table are the ones to compare against the leaderboard.

The top of the sub-100M field, using the leaderboard's published figures. tinctura-v1 is scored with the same formula but is not yet listed on the board.

ModelParamsIndex
Rose-1.5-Medium98.2M21.07
tinctura-v196.2M20.81
Surjo-100m97.7M18.86
cRia-LM-75M75.7M18.06
Rose-Medium97.8M17.73
Surjo-50M53.8M16.40

Against the leader the result is split. tinctura-v1 is ahead on ARC-Easy and PIQA, level on HellaSwag, and behind on ARC-Challenge and ArithMark-3. ArithMark-3 is most of the gap: 2.3 points there cost about 0.55 Index at its 0.65 weight, more than the whole 0.26 margin. Rose-1.5-Medium was trained on roughly 100B tokens according to its card, against 75B here.

What a third fewer parameters costs

Because tinctura-v1 follows cagliostro-v3 step for step, the two runs can be compared at identical points in training. Several of these comparisons were run side by side on the same machine with the same evaluation code, using checkpoints from both repositories. Every number in this section comes from the training-time harness, for both models.

Through the stable phase the smaller model trailed by about 1.7 Index over the first 20B tokens and by about 2.9 after that. Single evaluations carry roughly one point of noise, so individual gaps ranged from 0.5 to 3.9. Just before the cooldown, at step 646,000, the paired evaluation read 19.34 against 21.96.

The cooldown is where the two models part. cagliostro-v3 gained 4.6 Index over the last 15% of training. tinctura-v1 gained 1.4. The shortfall came in two places. From step 646,000 to 706,000 cagliostro-v3 added 1.95 and tinctura-v1 only 0.21. From 706,000 to 736,000 the two moved together, 1.32 and 1.23. Over the final stretch to the end cagliostro-v3 added another 1.32 while tinctura-v1 held flat. On ARC-Challenge and ArithMark-3 the smaller model finished the cooldown slightly lower than it started it.

After the first few billion tokens the training loss gap to cagliostro-v3 settled between 0.07 and 0.10 and stayed there through the cooldown. The sharp drop at 63.75B is the data mixture changing, not the model improving.

The cooldown was run twice

The first attempt at the cooldown hit a bug in the data loader. When training switches to the cooldown mixture, each process restarts its data-loading threads. Threads blocked on a full queue could miss the stop signal and keep producing batches from the old mixture next to the new ones, so each process trained on a blend. A standalone test of the loader leaked old-mixture batches in 9 of 10 trials, 20 to 44% of batches after the switch. The loss told on it: at the switch the loss fell by 0.26 where cagliostro-v3's fell by 0.49.

The loader was fixed so that every restart gets its own queue and old threads exit on their own, and the fixed version leaked nothing in the same test. The cooldown was then rerun from the last checkpoint before it, step 644,000. The rerun's loss fell by 0.45 at the switch and tracked cagliostro-v3 as expected. The weights in this repository come from the rerun.

The fix did not change the benchmarks in a measurable way. At step 706,000 the first attempt scored 19.68 and the rerun 19.55, which is within evaluation noise. The data-mixing bug was real, but it was not what held the cooldown gain down.

Model details

FieldValue
Parameters96,200,064
Non-embedding parameters78.2%
Layers18
Hidden size640
Intermediate size1,536
Attention heads10
Key/value heads5
AttentionGrouped query attention with cross-head subspace attenuation
ActivationSwiGLU
NormalizationRMSNorm, eps 1e-6
Positional encodingRoPE, theta 100,000
Context length2,048
Vocabulary32,768 BPE
EmbeddingsTied input and output
Logit cap15.0
Weightsfloat32 safetensors

The architecture is defined in this repository and is the same code as cagliostro-v3. trust_remote_code=True is required because CagliostroForCausalLM is not part of transformers.

Training data

The same mixtures and shards as cagliostro-v3. The first mixture covers the stable phase, and the second takes over when the cooldown begins at 85% of

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms