Model reference · open weights
Motif-3 is an open-weight language model from Motif-Technologies. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | Motif-Technologies |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 314.8B |
| Context | 256k tokens |
| Runs with | transformers |
| Based on | Motif-Technologies/Motif-3-Base |
| Released | 2026-08-07 |
| Popularity | 5k downloads / month |
| Licence | Open weights |
About
Motif 3 is a large-scale, decoder-only Mixture-of-Experts (MoE) language model with 314 billion total parameters and 13.2 billion parameters activated per token. It is built from the ground up by Motif Technologies following a fully in-house, proprietary design.
Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key–value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections (mHC), Expert-Specific PolyNorm activations, and a Multi-Token Prediction (MTP) head to improve optimization stability, expert specialization, and inference efficiency.
The model is pretrained on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora, with additional emphasis on Korean, reasoning-intensive, legal, and financial data. Post-training combines general supervised fine-tuning, six RL-trained specialist teachers, a software-engineering teacher, and Multi-teacher On-Policy Distillation (MOPD) into a single unified model.
For contextual comparison, Motif 3 is compared with strong open-weight models using scores reported on the corresponding benchmark leaderboards. All Motif 3 evaluations were performed with sampling temperature = 1.0, top-p = 0.95, and a maximum sequence length of 262,144 tokens. (*: public dataset only)
| Benchmark | Motif 3314B-A13B | MiniMax-3428B-A23B | GLM-5.1744B-A40B | Kimi-K2.61T-A32B | Qwen-3.7max | DS-v4-Pro1.6T-A49B |
|---|---|---|---|---|---|---|
| Agentic | ||||||
| GDPVal v2 | 38.7 | 44.4 | 37.8 | 34.4 | 39.0 | 40.2 |
| τ²-Bench Telecom | 94.7 | 88.9 | 97.7 | 95.9 | 94.7 | 96.2 |
| τ³-Banking | 35.3 | 15.3 | 13.6 | 23.3 | 12.0 | 30.1 |
| ITBench* | 51.5 | — | 40.3 | 31.2 | 42.5 | 38.3 |
| Coding | ||||||
| SWE-Bench Verified | 76.2 | 75.0 | 76.4 | 76.2 | 80.4 | 77.4 |
| Terminal-Bench 2.1 | 74.9 | 65.2 | 61.8 | 65.9 | 75.0 | 64.0 |
| SciCode | 40.6 | 45.4 | 43.8 | 53.5 | 53.5 | 50.0 |
| Reasoning & Knowledge | ||||||
| IMOAnswerBench | 83.2 | — | 83.8 | 81.8 | 90.0 | 89.8 |
| Apex-Shortlist | 75.5 | — | 71.1 | 77.4 | 44.5 | 85.8 |
| GPQA Diamond | 83.4 | 92.9 | 86.8 | 91.1 | 92.4 | 88.8 |
| HLE | 37.0 | 39.0 | 30.1 | 37.5 | 41.4 | 37.5 |
| CritPt | 6.6 | 3.7 | 4.6 | 8.0 | 11.4 | 12.9 |
| OmniScience — Accuracy | 30.1 | 16.7 | 23.7 | 32.6 | 31.0 | 42.9 |
| OmniScience — Non-Hallucination | 71.6 | 81.6 | 70.1 | 59.5 | 74 | 5.9 |
| Long Context & Instruction Following | ||||||
| AA-LCR | 72.3 | 80.3 | 68.0 | 76.7 | 75.0 | 70.0 |
| IFBench | 78.2 | 82.9 | 76.3 | 76.0 | 79.1 | 76.5 |
Motif 3 performs particularly well on agentic and tool-oriented benchmarks, while maintaining competitive performance across coding, mathematical reasoning, and general knowledge. On AA-Omniscience it pairs its accuracy with one of the highest non-hallucination scores, indicating a favorable balance between answering correctly and abstaining when unsupported.
[!NOTE] The architecture and distributed training framework used for Motif 3 are available at MotifTechnologies/motif3-training-example.
Motif 3 is a fully in-house design and introduces several custom components (full details in the technical report):
[!Note]
- Tested on B200 and H200 GPUs.
- The model ships with a built-in MTP head (
num_nextn_predict_layers=1), so it supports self-speculative decoding — add--speculative-configas shown below (num_speculative_tokens: 1is optimal for this model).- Supports online block-fp8 quantization with
--quantization modelopt_blockfp8- If you encounter any issues, please open an HF issue.
[!Tip] Looking for a smaller footprint? An NVFP4-quantized checkpoint is available at [Motif-Technologies/Motif-3-NVFP4](htt
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys motif-3 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (motif-3 below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"motif-3","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.