Pricing · one page, every price
Every price we charge is on this page. Per-token, per-image, pay-as-you-go — no subscription, no minimum, no hidden tiers. Where no grounded market comparison exists yet, the price is on request instead of invented.
Serverless Models
| Model | Unit | Price | Region | Status |
|---|---|---|---|---|
| Qwen3.8 27B — input | 1M tokens | €0.29 | eu-se-1 · eu-es-1 | Available · launch pricing |
| Qwen3.8 27B — output | 1M tokens | €1.77 | eu-se-1 · eu-es-1 | Available · launch pricing |
| Qwen3 Embedding | 1M tokens | €0.015 | eu-se-1 · eu-es-1 | Available |
| ERNIE Image Turbo | image | €0.02 | eu-se-1 · eu-es-1 | Available |
| FLUX.2 Klein 4B edits | image | €0.02 | eu-se-1 · eu-es-1 | Available |
| Whisper | minute | €0.005 | eu-se-1 · eu-es-1 | Available |
| Piper | 1M characters | €2.95 | eu-se-1 · eu-es-1 | Available |
| MiniMax Music 3 | track | €0.09 | eu-se-1 · eu-es-1 | Available |
Pay-as-you-go, no subscription, no minimum. Every API response reports token counts. Prices exclude VAT.
Managed GPU
| Hardware | From | Region | Status |
|---|---|---|---|
| NVIDIA DGX Spark (GB10) — 128 GB unified | €495 / month | eu-se-1 | Available |
| NVIDIA RTX 6000 Pro — 96 GB | €1.73 / GPU-hr | EU | On request |
| NVIDIA RTX 5090 ×4 — 32 GB each | €0.54 / GPU-hr | EU | On request |
| NVIDIA RTX 3090 ×4 — 24 GB each | €0.17 / GPU-hr | EU | On request |
| Model deployment — validated open models or your own weights | Request quote | EU | On request |
Deployment and operation included — machine, CUDA, runtime, endpoint. An engineer confirms the configuration and monthly price before anything is billed. Request deployment.
GPU VM
| GPU | Memory | Price | Region | Status |
|---|---|---|---|---|
| DGX Spark (GB10) | 128 GB unified | €495 / month | eu-se-1 | Available |
| RTX 6000 Pro | 96 GB | €1.73 / GPU-hr | EU | On request |
| RTX 5090 ×4 | 32 GB each | €0.54 / GPU-hr | EU | On request |
| RTX 3090 ×4 | 24 GB each | €0.17 / GPU-hr | EU | On request |
| H100 / H200 — against registered demand | 80 / 141 GB | — | EU | On request |
SSH access, dedicated machine, base environment included. Hardware without a price carries none until it lands. Deploy GPU.
The pricing rule
We price inference aggressively low: our target is 80% of the lowest current price from a proper inference supplier for the same model, tracked continuously. Found better pricing at a proper inference supplier? Tell us and we’ll look into lowering ours.
In practice: we maintain a price index of suppliers serving the same model and set our target at 80% of the lowest entry. A proper inference supplier is a provider operating and standing behind managed inference for the model — decentralized compute marketplaces don't count. The current Qwen3.8 27B prices (€0.29 / €1.77 per 1M tokens) were set against the index of 2026-08-26.
Why so cheap
Three reasons, none of them a trick:
| We own the hardware | Inference runs on NVIDIA DGX Spark systems AxForge owns and operates in Sweden — no cloud markup between you and the silicon. |
|---|---|
| Efficient models, right-sized machines | We serve efficient open-weight models on machines sized for exactly that job, so the hardware is actually utilized instead of idling expensively. |
| Real cost plus a thin margin | The price reflects what serving actually costs us, plus a thin margin. No VC-subsidized loss-leader pricing that has to snap back later. |
What the low price does not buy us: your data. Prompts and completions are processed in memory, never retained, never used to train — see the privacy policy.
FAQ
Not a public one today. API keys are allocated from the waiting list, and pay-as-you-go starts at the listed per-token prices — no subscription, no minimum.
These are the opening prices for the shared API, set by our pricing rule — 80% of the lowest current price from a proper inference supplier for the same model at the time we set them (index of 2026-08-26). It is not a teaser with a preset expiry: future prices follow the same rule, tracked against the market.
We track a price index of proper inference suppliers serving the same model and target 80% of the lowest current entry, tracked continuously. If you find better pricing at a proper inference supplier, tell us and we’ll look into lowering ours.
They can, in both directions, because the target tracks the market. What never changes with the price: EU residency, zero prompt retention, and no training on customer data.
Chat is billed per token with separate input and output rates; embeddings per input token; image generation per image. Every API response reports usage — including in streaming mode — so you can reconcile your bill exactly. List prices exclude VAT.
There are no published volume tiers. For sustained high volume, a dedicated system — a monthly DGX Spark rental serving only your traffic — is usually the better shape. Talk to an engineer and we’ll quote it.
Related