Model reference · open weights

Speck2

Available as managed deployment LLMs specklabs Text gen 1 variants 731 dl/mo

Speck2 is an open-weight language model from specklabs. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byspecklabs
TypeLanguage models
TaskText gen
Parameters (lead)141M
Context4k tokens
Runs withtransformers
Released2026-09-01
Popularity731 downloads / month
LicenceOpen weights

About

What Speck2 is

Speck2-140M is a 140.7M parameter English base language model that interleaves global grouped-query attention with gated causal convolution. It was pretrained from scratch on 20B tokens using a three-phase curriculum that shifts from broad web text toward math, synthetic, and stylistically diverse high-quality text.

This is a base model, not instruction-tuned or specialized in any way. It has no chat template and no safety alignment.

Read the full model card

Summary

PropertyValue
Parameters140,652,288
Training tokens20.0B
Training sequence length2,048
Configured max context4,096 (unvalidated beyond 2,048)
Vocabulary32,000 (Mistral v0.1 SentencePiece)
Release formatBF16 Safetensors
Validation loss / perplexity2.2403 / 9.396
CPU decode, batch 155.1 tok/s
RTX 3090 decode, batch 1247.3 tok/s

Architecture

18 residual blocks: 8 global attention + 10 gated causal convolution, each followed by a SwiGLU feed-forward.

ComponentValue
Hidden width768
Embedding width640
SwiGLU intermediate2,304
Attention heads (Q / KV)12 / 3
Head dimension64
Conv inner width384
Conv kernel sizes3, 5
RoPE theta10,000
RMSNorm epsilon1e-5

Input/output embeddings (640-wide) are tied and connect to the 768-wide residual stream via learned projections.

Usage

Speck2-140M works with the Transformers Auto classes through its bundled custom model and tokenizer code. Set trust_remote_code=True when loading it.

pip install "transformers==5.1.0" torch sentencepiece safetensors
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "specklabs/Speck2-140M"
device = "cuda" if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else "cpu"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype="auto",
).to(device)

prompt = "The meaning of life is"
inputs = tokenizer(prompt, return_tensors="pt").to(device)

output = model.generate(
    **inputs,
    max_new_tokens=64,
    do_sample=False,
)
generated = output[0, inputs.input_ids.shape[1] :]
print(tokenizer.decode(generated, skip_special_tokens=True))

The bundled generation path is validated for single-prompt greedy decoding. Direct forward passes support right-padded batches when use_cache=False.

Training

SettingValue
Optimizer steps305,176
Tokens per step65,536
Sequence length2,048
Peak LR1.5e-3 (cosine decay, 2,048-step warmup)
Weight decay0.1
Gradient clipping1.0
Training time41.40 hours
Estimated compute19.89 EFLOP

Muon optimized 2D matrix parameters; AdamW (beta 0.9/0.95, epsilon 1e-8) handled embeddings, norms, and conv kernels. The run used a single RTX 5090 rented through Vast.ai.

Training documents were globally deduplicated after NFKC, lowercase, and whitespace normalization. The curriculum used three mixture phases:

Token rangeEmphasis
0-14BBroad high-quality web, educational, synthetic, math, and encyclopedia text
14-18BIncreased FineMath, Cosmopedia, and Ultra-FineWeb-L3 multi-style text
18-20BStrongest concentration of synthetic and multi-style text, with continued math emphasis

Evaluation

The quality columns combine the Open SLM Leaderboard at revision 2eafcfc647b667e67f3b0288e9b67da497a78052 and BananaMind Base Bench 1.1 at revision d4aade51312889e8580963e1ce960c6eaef1a450. No chat template or generation was used for the four Speck evaluations.

Benchmarks and speed

ModelParamsTraining tokensOpen SLM Int IndexBananaMind Base Bench 1.1 EloCPU prefillCPU decodeRTX 3090 prefillRTX 3090 decodeBF16 memory @2KBF16 state @2K
BananaMind-2-Pro139M100B24.9611312,190 tok/s43.0 tok/s64,060 tok/s140.3 tok/s325.1 MiB60.0 MiB
SmolLM2-135M135M~2T27.1311192,201 tok/s47.4 tok/s64,814 tok/s157.7 tok/s301.6 MiB45.0 MiB
GPT-X2.5-135M135M75B25.1711062,042 tok/s47.2 tok/s55,346 tok/s125.0 tok/s302.6 MiB45.0 MiB
Supra2-100M-Base101M30B19.4110303,362 tok/s56.0 tok/s113,326 tok/s298.1 tok/s216.0 MiB24.0 MiB
Speck1-140M141M5B18.159652,252 tok/s55.1 tok/s74,323 tok/s247.3 tok/s281.3 MiB12.0 MiB
Speck1-140M-Instruct141M5B + 317M SFT17.7510012,285 tok/s55.3 tok/s73,398 tok/s246.7 tok/s280.3 MiB12.0 MiB
Speck1.1-140M-Instruct141M5B + 559M SFT17.9010022,315 tok/s56.9 tok/s74,941 tok/s243.6 tok/s280.3 MiB12.0 MiB
Speck2-140M141M20B20.019532,252 tok/s55.1 tok/s74,323 tok/s247.3 tok/s281.3 MiB12.0 MiB

Open SLM Int Index means the chance-normalized Intelligence Index reported by the Open SLM Leaderboard. BananaMind Base Bench 1.1 Elo means the overall Elo reported by BananaMind Base Bench 1.1. Speed and memory values are local batch-1 measurements described below. Reference models were pretrained on 1.5 to 100 times as many tokens. This is a parameter-adjacent comparison, not a compute-matched one.

Inference speed

Speed was measured locally at batch 1 with eager PyTorch, model-native caches, last-token logits, and tokenization excluded. Prefill uses 512 tokens. Decode measures 64 greedy cached steps after a 448-token prefix and includes argmax. CPU runs use FP32 with 16 threads; RTX 3090 runs use BF16. Reported th

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys speck2 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (speck2 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"speck2","messages":[{"role":"user","content":"Hello"}]}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms