Model reference · open weights

Tokle

NEW · this week LLMs techdotus Text gen 1 build Open weights 653 dl/mo

Tokle is an open-weight language model from techdotus. Tokle-3M (FP32) weighs 6 MB; the smallest configuration that runs it is RTX 3060 12 GB.

  • Tokle-3M is a decoder-only language model with 2.91M parameters designed for text generation tasks.
  • It supports a context length of 512 tokens and is trained exclusively on English data.
  • The model is released under the MIT license.

Summary of the techdotus/Tokle-3M model card, 2026-10-03

What it is

Released bytechdotus
Released2026-09-30
Parameters3M
VRAM6 MB for the weights

What it runs on

Memory and cards for Tokle-3M (FP32)

6 MBweights, file size
5 MBcache per 1K tokens
442 MBruntime overhead, at least
512 tokenscontext max
CardRequests at onceContext maxMemory
512 each
RTX 3060 12 GB1000+all 51211.6 GB
RTX 4060 Ti 16 GB1000+all 51215.4 GB
RTX 3090 24 GB1000+all 51223.4 GB
RTX 4090 24 GB1000+all 51223.4 GB
RTX 5090 32 GB1000+all 51231.0 GB
L40S 48 GB1000+all 51244.0 GB
A100 80 GB1000+all 51278.2 GB
H100 80 GB1000+all 51278.1 GB
RTX PRO 6000 Blackwell 96 GB1000+all 51293.8 GB
DGX Spark (GB10) 128 GB unified1000+all 512107 GB
H200 141 GB1000+all 512138 GB
B200 180 GB1000+all 512176 GB
Memory needed at each load
Requests at once512 tokens each
1450 MB
5461 MB
8469 MB
16490 MB
32533 MB
64617 MB

One card, with vLLM's small-card settings.

From the model card

What techdotus says about Tokle

Read the model card

Model Summary

Tokle-3M is a decoder-only language model with 2.91M parameters. It was first trained on 12B tokens with SPAB (Static Pairwise Attention Bias), a frozen table of 8.39M token-pair association scores built from Pointwise Mutual Information (PMI) over the training corpus, giving 11.3M parameters in total during this stage. During training, for every query-key pair, SPAB hashed the two token IDs into the table, retrieved their PMI value, scaled it by a learned per-head factor, and added it to the attention logits before softmax.

After this stage, the SPAB table was removed and the model was trained for an additional 0.5B tokens to distill the knowledge in the SPAB matrix into its own layers. As a result, Tokle-3M runs entirely on its 2.91M parameters at inference, with no SPAB table required.

Model Architecture

ParameterValue
ArchitectureDecoder-only transformer (RMSNorm, RoPE, GQA, SwiGLU)
Layers9
Hidden size (d_model)144
Attention heads3
KV heads (GQA)1 (multi-query attention)
Head dim48
FFN intermediate size432
Max sequence length512
Tie word embeddingsYes
PrecisionFP32 weights
Parameters2.91M

How to use

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "techdotus/Tokle-3M"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True).eval()

ids = tok("The climate change", return_tensors="pt")
with torch.no_grad():
    out = model.generate(**ids, max_new_tokens=32, do_sample=False,
                         repetition_penalty=1.3)  # greedy
print(tok.decode(out[0], skip_special_tokens=True))

Benchmark Results

All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology.

HellaSwagARC-EasyARC-ChallengePIQAArithMark-3
27.20%34.85%23.98%55.01%40.80%

Ablation: SPAB vs. Distilled

Stage 1 model (SPAB active) vs. the released Tokle-3M (SPAB removed and distilled).

ModelParamsInt IndexHellaSwagARC-EasyARC-ChalPIQAArithMark-3
Tokle-SPAB-3M11.3M (2.91M trainable + 8.39M frozen)9.1627.22%34.68%24.49%54.95%41.70%
Tokle-3M2.91M8.9227.20%34.85%23.98%55.01%40.80%

Comparison Results

All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology. Scores for the other models are from the Open SLM Leaderboard. Bold marks the best result in each column.

ModelParamsInt IndexHellaSwagARC-EasyARC-ChalPIQAArithMark-3
Tokle-3M (Tech.us)2.91M8.9227.20%34.85%23.98%55.01%40.80%
Ember-2 (SurjoLabs)2.96M×27.2127.28%33.42%22.01%55.11%35.90%
BananaMind-2-Micro (BananaMind)2.9M6.0128.27%33.12%21.93%53.21%34.00%
GPT-S-1.4M (Axiomic Labs)1.4M5.4026.89%31.57%21.93%55.17%30.20%

Training Details

Tokle-3M was trained in two stages on the same data mixture.

StageTokensSPABParameters
1. Pretraining12BActive (frozen PMI table)11.3M (2.91M trainable + 8.39M frozen)
2. Distillation0.5BRemoved2.91M

Stage 2 lets the trained weights absorb the prior the SPAB table had been providing, so the released model is self-contained rather than losing that knowledge when the table is removed.

Training Data

We trained on a curated mixture with a strict cleaning pipeline that also removed topics not useful for a model of this size.

SourcePercentage
FineWeb-Edu43.1%
Cosmopedia24.3%
OpenMathInstruct-213.5%
Tiny Strange Textbooks9.0%
MegaScience (medicine & biology, custom curated)5.0%
High-Quality English Sentences3.0%
ScienceQA1.2%
Orca-Math Word Problems 200k0.9%
Total100%
  • Tokenizer: all data was tokenized with the model's 5,048-token BPE tokenizer, and 1% was held out for validation.
  • Blending: sources were blended per dataset using the weights above.

Limitations

  • Tiny model: with 2.91M parameters and 144-dim hidden states, generations are often repetitive, incoherent or factually wrong. The model is a research artifact for studying small-scale LMs, not an assistant.
  • Short context: 512 tokens maximum. RoPE tables are not built beyond that length.
  • English only: trained on English web, educational, synthetic and math text.
  • Not instruction-tuned or safety-aligned: it may reproduce biases present in web data.

License

Model weights and code: MIT.

Citation

@misc{tokle2026,
  title        = {{Tokle-3M}: Pointwise Mutual Information as a Removable
                  Inductive Bias for Self-Attention},
  author       = {{Tech.us Team}},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/techdotus/Tokle-3M}}
}

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms