Model reference · open weights
Supra is an open-weight language model from SupraLabs. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | SupraLabs |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 52M |
| Context | 1k tokens |
| Runs with | transformers |
| Released | 2026-05-21 |
| Popularity | 1k downloads / month |
| Licence | Open weights |
About
Supra-50M is a compact 50M-parameter BASE causal language model built by SupraLabs, trained from scratch using a Llama-style architecture on 20 billion tokens of high-quality educational web text. Despite being significantly smaller than comparable open models, it achieves competitive or superior results on several key benchmarks. It's our first SupraLabs Scaling Up Plan model.
| Benchmark | Supra-50M (ours) | GPT-2 (124M) | SmolLM-135M | OpenELM-270M |
|---|---|---|---|---|
| Parameters | 50M | 124M (2.5×) | 135M (2.7×) | 270M (5.4×) |
| BLiMP (linguistics) | 76.3% | 63.0% | 69.8% | (N/A) |
| SciQ (science) | 77.2% | 53.2% | 73.4% | 84.70% |
| ARC-Easy (knowledge) | 52.2% | 42.0% | 49.2% | 45.08% |
| PIQA (logic) | 62.2% | 63.0% | 67.3% | 69.75% |
| HellaSwag (context) | 31.8% | 29.5% | 42.0% | 46.71% |
| Task | Metric | Value |
|---|---|---|
| arc_easy | acc,none | 0.5185 |
| arc_easy | acc_stderr,none | 0.0103 |
| arc_easy | acc_norm,none | 0.4600 |
| arc_easy | acc_norm_stderr,none | 0.0102 |
| arc_challenge | acc,none | 0.2159 |
| arc_challenge | acc_stderr,none | 0.0120 |
| arc_challenge | acc_norm,none | 0.2517 |
| arc_challenge | acc_norm_stderr,none | 0.0127 |
| hellaswag | acc,none | 0.2903 |
| hellaswag | acc_stderr,none | 0.0045 |
| hellaswag | acc_norm,none | 0.3172 |
| hellaswag | acc_norm_stderr,none | 0.0046 |
| winogrande | acc,none | 0.5154 |
| winogrande | acc_stderr,none | 0.0140 |
| piqa | acc,none | 0.6251 |
| piqa | acc_stderr,none | 0.0113 |
| piqa | acc_norm,none | 0.6219 |
| piqa | acc_norm_stderr,none | 0.0113 |
| openbookqa | acc,none | 0.1860 |
| openbookqa | acc_stderr,none | 0.0174 |
| openbookqa | acc_norm,none | 0.3080 |
| openbookqa | acc_norm_stderr,none | 0.0207 |
| boolq | acc,none | 0.5303 |
| boolq | acc_stderr,none | 0.0087 |
Supra-50M is based on the LlamaForCausalLM architecture with the following configuration:
| Hyperparameter | Value |
|---|---|
| Architecture | Llama (decoder-only transformer) |
| Parameters | ~50M |
vocab_size | 32,000 |
hidden_size | 512 |
intermediate_size | 1,408 |
num_hidden_layers | 12 |
num_attention_heads | 8 |
num_key_value_heads | 4 (GQA) |
max_position_embeddings | 1,024 |
rope_theta | 10,000 |
tie_word_embeddings | True |
| Property | Value |
|---|---|
| Dataset | HuggingFaceFW/fineweb-edu (sample-100BT split) |
| Total tokens | 20,000,000,000 (20B) |
| Sequence length | 1,024 tokens |
| Storage format | Memory-mapped binary (uint16, ~40 GB) |
A custom Byte-Level BPE tokenizer was trained from scratch on 500,000 documents sampled from fineweb-edu (sample-10BT).
| Property | Value |
|---|---|
| Type | ByteLevelBPETokenizer |
| Vocabulary size | 32,000 |
| Min frequency | 2 |
| Special tokens | , , , , `` |
| Parameter | Value |
|---|---|
| Epochs | 1 |
| Per-device batch size | 32 |
| Gradient accumulation steps | 4 |
| Effective batch size | 128 × 1,024 tokens |
| Learning rate | 6e-4 |
| LR scheduler | Cosine |
| Warmup ratio | 2% |
| Optimizer | AdamW Fused (adam_beta1=0.9, adam_beta2=0.95) |
| Weight decay | 0.1 |
| Max grad norm | 1.0 |
| Precision | bfloat16 |
torch.compile | Enabled |
| Hardware | Single GPU |
| Final loss | 3.259 |
from transformers import pipeline
import torch
print("[*] Loading Supra-50M model from Hugging Face Hub...")
pipe = pipeline(
"text-generation",
model="SupraLabs/Supra-50M_BASE",
device_map="auto",
torch_dtype=torch.float16 if torch.cuda.is_available() else torch.float32
)
def generate_text(prompt, max_new_tokens=150):
result = pipe(
prompt,
max_new_tokens=max_new_tokens,
do_sample=True,
temperature=0.5,
top_k=25,
top_p=0.9,
repetition_penalty=1.2,
pad_token_id=pipe.tokenizer.pad_token_id,
eos_token_id=pipe.tokenizer.eos_token_id
)
return result[0]['generated_text']
# Example
prompt = "The importance of education is"
print(f"\nPrompt: {prompt}")
print("-" * 40)
print("\nOutput:\n" + generate_text(prompt))
Prompt: "The main concept of physics is "
The main concept of physics is iffy, and the idea that we can make things behave in a certain way. The most important part of physics is called quantum mechanics which states that all particles are made up of energy (energy) and matter (matter). In physics, there are two types of particles: elementary particles and exotic ones. These particles have properties like mass, speed or momentum but they don’t interact with each other to form new objects. This is because these particles do not exist independently from one another. In this case, an exotic particle might be created by adding more energy into its structure than it would take for a normal particle. However, when you add additional energy to an exotic particle, the new object will become smaller and larger until it becomes too large to fit within the existing structure. If you think about how light travels through space, it takes around 20 billion years before the light reaches our eyes. Light waves travel faster than light at high speeds so if we could create some kind of light wave, then we wouldn’t need any special equipment. It just needs a few hundred millionths of a second to produce light rays. So even though the light is moving along the same path as the current, the speed of light is different depending on where the light hits the
Prompt: "Artificial intelligence is "
Artificial intelligence is iffy, it can be used to make intelligent machines that could take over the world. What does Artificial Intelligence mean? AI refers to ar
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys supra for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (supra below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"supra","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.