Models & pricing
One key serves every model below. Pass the API model name in the
model field of the matching endpoint. Prices are per-use, in
EUR, excluding VAT.
Served on the shared API today
| Model | Endpoint | API model name | Context | Price |
|---|---|---|---|---|
| Qwen3.8 27B | /v1/chat/completions | qwen3.8-27b-nvfp4 | 65,536 | €0.29 / 1M in · €1.77 / 1M out |
| Qwen3 Embedding | /v1/embeddings | qwen3-embed | 32,768 | €0.015 / 1M |
| ERNIE Image Turbo | /v1/images/generations | ernie-image-turbo | — | €0.02 / image |
| FLUX.2 Klein 4B | /v1/images/edits | flux2-klein-4b | — | €0.02 / image |
| Whisper | /v1/audio/transcriptions | whisper | — | €0.005 / minute |
| Piper | /v1/audio/speech | piper | — | €2.95 / 1M characters |
| MiniMax Music 3 | /v1/audio/music | minimax-music3 | — | €0.09 / track |
Chat prices are launch pricing. Each linked page carries the model's specs, measured performance and data-handling details; the API pages in these docs (chat, embeddings, images, audio) carry the request and response shapes.
How we price
We price inference aggressively low: our target is 80% of the lowest current price from a proper inference supplier for the same model, tracked continuously. Found better pricing at a proper inference supplier? Tell us and we'll look into lowering ours.
Where the tracked market has no comparable listing — Whisper, Piper, image editing, music — prices are launch prices anchored to the closest public list price. Committed volume: talk to an engineer.
On request
These models are validated on AxForge hardware and can be brought hot on a dedicated system for you. They are not on the shared API today — the page for each says so plainly.
| Model | Type | Context |
|---|---|---|
| Qwen3.6 35B A3B | LLM (MoE, 3B active) | 65,536 |
| Qwen3 30B A3B | LLM (MoE) | 65,536 |
| Gemma-4 26B | LLM, vision-capable | 262,144 |
| Mistral Small 3.2 | LLM | 32,768 |
| SDXL | Image generation | — |
| Qwen-Image | Image generation & editing | — |
Want one of these hot, or a different open model on dedicated hardware? Talk to an engineer.