Model reference · open weights
Tokle is an open-weight language model from techdotus. Tokle-3M (FP32) weighs 6 MB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the techdotus/Tokle-3M model card, 2026-10-03
What it is
| Released by | techdotus |
|---|---|
| Released | 2026-09-30 |
| Parameters | 3M |
| VRAM | 6 MB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory |
|---|---|---|---|
| 512 each | |||
| RTX 3060 12 GB | 1000+ | all 512 | 11.6 GB |
| RTX 4060 Ti 16 GB | 1000+ | all 512 | 15.4 GB |
| RTX 3090 24 GB | 1000+ | all 512 | 23.4 GB |
| RTX 4090 24 GB | 1000+ | all 512 | 23.4 GB |
| RTX 5090 32 GB | 1000+ | all 512 | 31.0 GB |
| L40S 48 GB | 1000+ | all 512 | 44.0 GB |
| A100 80 GB | 1000+ | all 512 | 78.2 GB |
| H100 80 GB | 1000+ | all 512 | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 1000+ | all 512 | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 1000+ | all 512 | 107 GB |
| H200 141 GB | 1000+ | all 512 | 138 GB |
| B200 180 GB | 1000+ | all 512 | 176 GB |
| Requests at once | 512 tokens each |
|---|---|
| 1 | 450 MB |
| 5 | 461 MB |
| 8 | 469 MB |
| 16 | 490 MB |
| 32 | 533 MB |
| 64 | 617 MB |
One card, with vLLM's small-card settings.
From the model card
Tokle-3M is a decoder-only language model with 2.91M parameters. It was first trained on 12B tokens with SPAB (Static Pairwise Attention Bias), a frozen table of 8.39M token-pair association scores built from Pointwise Mutual Information (PMI) over the training corpus, giving 11.3M parameters in total during this stage. During training, for every query-key pair, SPAB hashed the two token IDs into the table, retrieved their PMI value, scaled it by a learned per-head factor, and added it to the attention logits before softmax.
After this stage, the SPAB table was removed and the model was trained for an additional 0.5B tokens to distill the knowledge in the SPAB matrix into its own layers. As a result, Tokle-3M runs entirely on its 2.91M parameters at inference, with no SPAB table required.
| Parameter | Value |
|---|---|
| Architecture | Decoder-only transformer (RMSNorm, RoPE, GQA, SwiGLU) |
| Layers | 9 |
| Hidden size (d_model) | 144 |
| Attention heads | 3 |
| KV heads (GQA) | 1 (multi-query attention) |
| Head dim | 48 |
| FFN intermediate size | 432 |
| Max sequence length | 512 |
| Tie word embeddings | Yes |
| Precision | FP32 weights |
| Parameters | 2.91M |
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "techdotus/Tokle-3M"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True).eval()
ids = tok("The climate change", return_tensors="pt")
with torch.no_grad():
out = model.generate(**ids, max_new_tokens=32, do_sample=False,
repetition_penalty=1.3) # greedy
print(tok.decode(out[0], skip_special_tokens=True))
All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology.
| HellaSwag | ARC-Easy | ARC-Challenge | PIQA | ArithMark-3 |
|---|---|---|---|---|
| 27.20% | 34.85% | 23.98% | 55.01% | 40.80% |
Stage 1 model (SPAB active) vs. the released Tokle-3M (SPAB removed and distilled).
| Model | Params | Int Index | HellaSwag | ARC-Easy | ARC-Chal | PIQA | ArithMark-3 |
|---|---|---|---|---|---|---|---|
| Tokle-SPAB-3M | 11.3M (2.91M trainable + 8.39M frozen) | 9.16 | 27.22% | 34.68% | 24.49% | 54.95% | 41.70% |
| Tokle-3M | 2.91M | 8.92 | 27.20% | 34.85% | 23.98% | 55.01% | 40.80% |
All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology. Scores for the other models are from the Open SLM Leaderboard. Bold marks the best result in each column.
| Model | Params | Int Index | HellaSwag | ARC-Easy | ARC-Chal | PIQA | ArithMark-3 |
|---|---|---|---|---|---|---|---|
| Tokle-3M (Tech.us) | 2.91M | 8.92 | 27.20% | 34.85% | 23.98% | 55.01% | 40.80% |
| Ember-2 (SurjoLabs) | 2.96M×2 | 7.21 | 27.28% | 33.42% | 22.01% | 55.11% | 35.90% |
| BananaMind-2-Micro (BananaMind) | 2.9M | 6.01 | 28.27% | 33.12% | 21.93% | 53.21% | 34.00% |
| GPT-S-1.4M (Axiomic Labs) | 1.4M | 5.40 | 26.89% | 31.57% | 21.93% | 55.17% | 30.20% |
Tokle-3M was trained in two stages on the same data mixture.
| Stage | Tokens | SPAB | Parameters |
|---|---|---|---|
| 1. Pretraining | 12B | Active (frozen PMI table) | 11.3M (2.91M trainable + 8.39M frozen) |
| 2. Distillation | 0.5B | Removed | 2.91M |
Stage 2 lets the trained weights absorb the prior the SPAB table had been providing, so the released model is self-contained rather than losing that knowledge when the table is removed.
We trained on a curated mixture with a strict cleaning pipeline that also removed topics not useful for a model of this size.
| Source | Percentage |
|---|---|
| FineWeb-Edu | 43.1% |
| Cosmopedia | 24.3% |
| OpenMathInstruct-2 | 13.5% |
| Tiny Strange Textbooks | 9.0% |
| MegaScience (medicine & biology, custom curated) | 5.0% |
| High-Quality English Sentences | 3.0% |
| ScienceQA | 1.2% |
| Orca-Math Word Problems 200k | 0.9% |
| Total | 100% |
Model weights and code: MIT.
@misc{tokle2026,
title = {{Tokle-3M}: Pointwise Mutual Information as a Removable
Inductive Bias for Self-Attention},
author = {{Tech.us Team}},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/techdotus/Tokle-3M}}
}
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.