Model reference · open weights
SuperApriel is an open-weight language model from ServiceNow-AI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Maker | ServiceNow-AI |
|---|---|
| Type | Language models |
| Task | Text gen |
| Runs with | transformers |
| Released | 2026-04-07 |
| Popularity | 126 downloads / month |
| Licence | Open weights |
About
A 15B-parameter token-mixer supernet with 8 optimized deployment presets spanning 1.0× to 10.7× decode throughput at 32K sequence length, all from a single checkpoint. Derived from Apriel-1.6 through stochastic distillation and targeted supervised fine-tuning.
See the report for detailed benchmarks, quality retention curves, and the full story.
Each preset is a specific assignment of one mixer per layer. All presets share the same checkpoint—only mixer selection differs at inference time.
| Preset | FA | SWA | KDA | GDN | Speedup @32k | Speedup @16k | Avg Acc. | Quality Retention |
|---|---|---|---|---|---|---|---|---|
| all-attention | 48 | 0 | 0 | 0 | 1.0× | 1.0× | 74.2 | 100% |
| Reg|Lklhd‑26 | 12 | 26 | 6 | 4 | 2.85× | 1.5× | 71.1 | 96% |
| Idealized|All‑18 | 13 | 32 | 1 | 2 | 1.99× | 1.1× | 71.8 | 97% |
| Reg|Lklhd‑18 | 3 | 25 | 4 | 16 | 4.76× | 2.2× | 69.7 | 94% |
| Idealized|Lklhd‑6 | 0 | 30 | 5 | 13 | 6.2× | 2.4× | 66.8 | 90% |
| Idealized|All‑6 | 0 | 30 | 5 | 13 | 6.13× | 2.5× | 65.3 | 88% |
| Reg|Lklhd‑13 | 0 | 16 | 13 | 19 | 6.9× | 2.7× | 60.2 | 81% |
| Reg|Lklhd‑10 | 0 | 10 | 5 | 33 | 10.69× | 4.2× | 57.2 | 77% |
| Config | @32k | AIME'24 | AIME'25 | MATH-500 | GSM8K | FDA | SWDE | RULER | Tau2 | MMLU-Pro | AIME(NV) | GPQA | HLE | LCB | IFBench | All |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| all-attention | 1.0× | 93.3 | 86.7 | 91.8 | 92.3 | 78.3 | 89.5 | 79.4 | 56.7 | 76.8 | 82.7 | 72.0 | 8.2 | 68.6 | 63.1 | 74.2 |
| Reg|Lklhd‑26 | 2.85× | 86.7 | 83.3 | 92.0 | 91.1 | 79.9 | 88.2 | 74.4 | 30.7 | 76.2 | 80.0 | 70.7 | 10.0 | 69.2 | 63.1 | 71.1 |
| Idealized|All‑18 | 1.99× | 90.0 | 86.7 | 92.0 | 92.1 | 78.0 | 86.6 | 67.1 | 52.6 | 76.3 | 82.2 | 68.2 | 6.9 | 67.3 | 58.8 | 71.8 |
| Reg|Lklhd‑18 | 4.76× | 86.7 | 76.7 | 92.4 | 91.7 | 81.7 | 89.8 | 60.5 | 46.2 | 76.3 | 74.4 | 68.7 | 6.6 | 64.8 | 59.3 | 69.7 |
| Idealized|Lklhd‑6 | 6.2× | 83.3 | 76.7 | 92.4 | 92.3 | 76.9 | 88.9 | 66.1 | 40.4 | 73.6 | 62.2 | 65.0 | 6.1 | 57.0 | 54.9 | 66.8 |
| Idealized|All‑6 | 6.13× | 83.3 | 80.0 | 92.2 | 91.7 | 75.6 | 87.4 | 61.9 | 34.2 | 73.3 | 56.7 | 61.0 | 5.9 | 55.9 | 55.3 | 65.3 |
| Reg|Lklhd‑13 | 6.9× | 76.7 | 73.3 | 90.4 | 91.2 | 68.6 | 85.1 | 57.0 | 28.6 | 69.3 | 26.7 | 61.2 | 5.5 | 52.6 | 57.1 | 60.2 |
| Reg|Lklhd‑10 | 10.69× | 76.7 | 66.7 | 90.6 | 90.8 | 65.2 | 82.9 | 48.6 | 23.4 | 68.2 | 24.4 | 52.5 | 4.5 | 50.2 | 56.2 | 57.2 |
| Model | Speedup @32k | Math (Avg) | All Tasks |
|---|---|---|---|
| Super Apriel all-attention | 1.0× | 91.0 | 74.2 |
| Super Apriel Reg|Lklhd‑26 | 2.85× | 88.3 | 71.1 |
| Super Apriel Reg|Lklhd‑18 | 4.76× | 86.8 | 69.7 |
| Super Apriel Idealized|Lklhd‑6 | 6.2× | 86.2 | 66.8 |
| Apriel-H1 15B | 1.97× | 80.4 | 58.4 |
| Nemotron-Nano 12B v2 | 5.85× | 74.5 | 62.4 |
| Falcon-H1R 7B | 4.61× | 78.6 | 64.9 |
| Nemotron-3-Nano 30B | 4.09× | 89.0 | 72.6 |
SuperApriel-15b-Instruct is trained in two stages:
Stage 1 — Stochastic Distillation: All four mixer types trained simultaneously via distillation from frozen Apriel-1.6 teacher on 266B tokens. See SuperApriel-15b-Base.
Stage 2 — Targeted SFT: Supervised fine-tuning on 60B tokens with 8 Pareto-optimal placements identified via Bayesian placement optimization. Shared parameters (FFNs, embeddings, norms) remain frozen; only mixer weights are trained.
| Component | Details |
|---|---|
| Parameters | 15B |
| Decoder layers | 48 |
| Query / KV heads | 32 / 8 (grouped-query attention), d_h = 128 |
| Hidden dimension | 5,120 |
| FFN width | 14,336 (SiLU-gated) |
| Vocabulary | 131,072 tokens |
| Vision encoder | Pixtral (16×16 patches) |
| Mixer | Time | Memory | Description |
|---|---|---|---|
| Full Attention (FA) | O(n²) | O(n) KV cache | Standard grouped-query attention |
| Sliding Window (SWA) | O(w·n) | O(w) | Local window of 4,096 tokens |
| Gated DeltaNet (GDN) | O(n) | O(1) fixed state | Matrix-valued recurrent state with delta rule |
| Kimi Delta Attention (KDA) | O(n) | O(1) fixed state | Linear attention with channel-wise gating |
The recommended serving backend is vLLM with the Fast-LLM plugin, which supports preset selection and runtime switching. For simpler use cases, Transformers is also supported (see Use with Transformers below).
Preset selection and throughput-optimized serving require the vLLM plugin from Fast-LLM. Two serving modes are available:
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys superapriel for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (superapriel below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"superapriel","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.