Model reference · open weights
Tokle-SPAB is an open-weight language model from techdotus. Tokle-SPAB-3M (FP32) weighs 23 MB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the techdotus/Tokle-SPAB-3M model card, 2026-10-03
What it is
| Released by | techdotus |
|---|---|
| Released | 2026-09-29 |
| Parameters | 3M |
| VRAM | 23 MB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory |
|---|---|---|---|
| 512 each | |||
| RTX 3060 12 GB | 1000+ | all 512 | 11.6 GB |
| RTX 4060 Ti 16 GB | 1000+ | all 512 | 15.4 GB |
| RTX 3090 24 GB | 1000+ | all 512 | 23.4 GB |
| RTX 4090 24 GB | 1000+ | all 512 | 23.4 GB |
| RTX 5090 32 GB | 1000+ | all 512 | 31.0 GB |
| L40S 48 GB | 1000+ | all 512 | 44.0 GB |
| A100 80 GB | 1000+ | all 512 | 78.2 GB |
| H100 80 GB | 1000+ | all 512 | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 1000+ | all 512 | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 1000+ | all 512 | 107 GB |
| H200 141 GB | 1000+ | all 512 | 138 GB |
| B200 180 GB | 1000+ | all 512 | 176 GB |
| Requests at once | 512 tokens each |
|---|---|
| 1 | 467 MB |
| 5 | 478 MB |
| 8 | 486 MB |
| 16 | 507 MB |
| 32 | 549 MB |
| 64 | 634 MB |
One card, with vLLM's small-card settings.
From the model card
Tokle-SPAB-3M is a decoder-only language model trained on 12B tokens, with 2.91M trainable parameters and 8.39M frozen SPAB parameters, for a total of 11.3M parameters. Its main architectural addition is SPAB (Static Pairwise Attention Bias), a frozen table of token-pair association scores built from Pointwise Mutual Information (PMI) over the training corpus and added to the attention logits.
For every query-key pair, SPAB hashes the two token IDs into the table, pulls out their PMI value, multiplies it by a learned per-head scale, and adds it to the attention logits before softmax. The bias ignores position and depends only on which tokens are involved, so the model starts training already knowing which tokens tend to co-occur. It only has to learn how much to trust that prior.
| Parameter | Value |
|---|---|
| Architecture | Custom decoder-only transformer + SPAB |
| Layers | 9 |
| Hidden size (d_model) | 144 |
| Attention heads | 3 |
| KV heads (GQA) | 1 (multi-query attention) |
| Head dim | 48 |
| FFN intermediate size | 432 |
| Max sequence length | 512 |
| Trainable parameters | 2,908,947 |
| Frozen SPAB table | 8,388,608 (float32 buffer) |
This model uses a custom architecture, so it needs trust_remote_code=True.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "techdotus/Tokle-SPAB-3M"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True).eval()
ids = tok("The climate change", return_tensors="pt")
with torch.no_grad():
out = model.generate(**ids, max_new_tokens=32, do_sample=False, repetition_penalty=1.3) # greedy
print(tok.decode(out[0], skip_special_tokens=True))
All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology:
| Hellaswag | ARC-Easy | ARC-Challenge | PIQA | Arithmark-3 |
|---|---|---|---|---|
| 27.22% | 34.68% | 24.49% | 54.95% | 41.70% |
We trained on a curated mixture with a strict cleaning pipeline that also removed topics not useful for a model of this size.
| Source | Percentage |
|---|---|
| FineWeb-Edu | 43.1% |
| Cosmopedia | 24.3% |
| OpenMathInstruct-2 | 13.5% |
| Tiny Strange Textbooks | 9.0% |
| MegaScience (medicine & biology, custom curated) | 5.0% |
| High-Quality English Sentences | 3.0% |
| ScienceQA | 1.2% |
| Orca-Math Word Problems 200k | 0.9% |
| Total | 100% |
Model weights and code: MIT.
@misc{tokle2026,
title = {{Tokle-SPAB-3M}: Pointwise Mutual Information as an Inductive Bias for Self-Attention},
author = {{Tech.us Team}},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/techdotus/Tokle-SPAB-3M}}
}
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.