Model reference · open weights
Puro is an open-weight language model from thu-pacman. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | thu-pacman |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 2.0B |
| Context | 4k tokens |
| Runs with | transformers |
| Released | 2026-08-16 |
| Popularity | 1k downloads / month |
| Licence | Open weights |
About
Under our fixed 15-benchmark base-model evaluation, one checkpoint in the Puro-2B collection beats Qwen2-1.5B at about $4.4K; the canonical final model goes further, approaching Qwen2.5-1.5B at a measured rental-equivalent accelerator cost of $6,891.
How Far Can a Poor Lab Go with RTX 5090s?
Puro-2B (普罗-2B) is a 2B-parameter dense causal language model pretrained from scratch on 1.4T tokens. It uses a Qwen3-1.7B-compatible architecture with untied input and output embeddings, blockwise FP8 training, the MuonH optimizer, and a two-phase data recipe. Training ran entirely on consumer-grade NVIDIA RTX 5090 GPUs.
The architecture is based on the Qwen3-1.7B configuration, not on pretrained Qwen weights. Puro-2B starts from random initialization.
Puro-2B is intended to make billion-parameter pretraining inspectable and affordable for smaller research groups. The release covers more than the final weights:
The main recipe combines RTX 5090 infrastructure, blockwise FP8, MuonH with hyperball constraints, proxy-guided data selection, and a curriculum-aware late continuation followed by checkpoint averaging.
The collection contains multiple checkpoints with different Phase 2 budgets and recipes. The report's approximately $4.4K result is an observed uniform-recipe checkpoint that already exceeds Qwen2-1.5B on the report's 15-task aggregate. It is not the canonical final checkpoint.
The canonical Puro-2B-Base model is the strongest released endpoint. Its
production run used 22,514 measured active-training GPU-hours, corresponding
to $6,891 under the report's normalized RTX 5090 rental rate.
These figures are accelerator-only reproduction estimates. They exclude data acquisition and preprocessing, proxy and ablation experiments, failed runs, post-training, evaluation, storage, networking, and research labor. They should not be read as the total cost of developing the project.
The scaling-law panel labels points by cumulative reproduction cost. The model catalog below maps those costs to Phase 2 budget fractions. Each fraction applies only to Phase 2 data exposure, while the cost includes the shared Phase 1 run.
| Property | Value |
|---|---|
| Model type | Dense decoder-only causal language model |
| Parameters | Approximately 2B |
| Initialization | From scratch |
| Architecture | Qwen3-1.7B configuration with untied embeddings |
| Hidden size | 2,048 |
| Transformer layers | 28 |
| Attention heads / KV heads | 16 / 8 |
| Feed-forward size | 6,144 |
| Vocabulary size | 151,936 |
| Context length | 4,096 tokens |
| Export class | Qwen3ForCausalLM |
| Weight format | Safetensors |
This is a pretrained base model. It has not been instruction-tuned or preference-aligned and should not be expected to behave like a chat assistant.
Use a Transformers release that supports the Qwen3 configuration:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "thu-pacman/Puro-2B-Base"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
prompt = "The central limit theorem states that"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=128,
do_sample=True,
temperature=0.7,
top_p=0.9,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
All numbers below come from the same deterministic OpenCompass pipeline in the technical report. The comparison uses pretrained/base checkpoints throughout. Generation tasks use greedy decoding; multiple-choice tasks use fixed token-likelihood ranking. Scores are percentages.
| Model | Math + Code (4) | Reasoning + Knowledge (11) | Overall (15) |
|---|---|---|---|
| Qwen2-1.5B | 40.29 | 60.54 | 55.14 |
| Puro-2B | 43.50 | 63.02 | 57.81 |
| Qwen2.5-1.5B | 47.52 | 65.53 | 60.73 |
The four math and code tasks are GSM8K, MATH, sanitized-MBPP, and HumanEval. The eleven reasoning and knowledge tasks are MMLU, MMLU-Pro, ARC-Challenge, ARC-Easy, BoolQ, CommonsenseQA, HellaSwag, PIQA, SocialIQA, WinoGrande, and BBH. Each displayed average is an unweighted arithmetic mean.
The Puro Cost Scaling Law fits five single-run Phase 2 uniform-budget points. It is a recipe-specific empirical scale-down relationship, not a universal law. The fit has no uncertainty interval, and the available experiments do not isolate curriculum ordering, constant-LR continuation, and checkpoint averaging as independent causal gains.
| Setting | Phase 1 | Phase 2 |
|---|---|---|
| Tokens consumed | 439B | 960B |
| RTX 5090 GPUs | 24 | 96 |
| Parallelism (TP / PP / DP) | 1 / 2 / 12 | 1 / 4 / 24 |
| Base learning rate | 5.00e-3 -> 1.04e-3 | 1.04e-3 -> 1.00e-5 |
| Schedule | Power decay | Linear decay, then selected constant-LR continuation |
| Median TFLOP/s/GPU | 238 | 192 |
Both phases use a sequence length of 4,096, a global batch size of 1,536 sequences, and a micro-batch size of 2. Main Transformer linear-layer GEMMs use blockwise E4M3 FP8; numerically sensitive operations, master weights, and optimizer states remain in BF16 or FP32 as appropriate.
Selected approximately scale-invariant matrix weights are updated by MuonH with hyperball projection and zero weight decay. The remaining parameters use AdamW with weight decay 0.1. The MuonH matrix group applies a 10x multiplier to the shared base l
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys puro for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (puro below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"puro","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.