Pricing · one page, every price

Pricing

Every price we charge is on this page. Per-token, per-image, pay-as-you-go — no subscription, no minimum, no hidden tiers. Where no grounded market comparison exists yet, the price is on request instead of invented.

€0.29 / 1M input · €1.77 / 1M output — Qwen3.8 27B eu-se-1 · Stockholm eu-es-1 · Málaga
Get an API key Talk to an engineer Launch pricing. Keys are allocated from the waiting list. Prices exclude VAT.

Serverless Models

Pay per use, in EUR

ModelUnitPriceRegionStatus
Qwen3.8 27B — input1M tokens€0.29eu-se-1 · eu-es-1Available · launch pricing
Qwen3.8 27B — output1M tokens€1.77eu-se-1 · eu-es-1Available · launch pricing
Qwen3 Embedding1M tokens€0.015eu-se-1 · eu-es-1Available
ERNIE Image Turboimage€0.02eu-se-1 · eu-es-1Available
FLUX.2 Klein 4B editsimage€0.02eu-se-1 · eu-es-1Available
Whisperminute€0.005eu-se-1 · eu-es-1Available
Piper1M characters€2.95eu-se-1 · eu-es-1Available
MiniMax Music 3track€0.09eu-se-1 · eu-es-1Available

Pay-as-you-go, no subscription, no minimum. Every API response reports token counts. Prices exclude VAT.

Managed GPU

We deploy and operate it

HardwareFromRegionStatus
NVIDIA DGX Spark (GB10) — 128 GB unified€495 / montheu-se-1Available
NVIDIA RTX 6000 Pro — 96 GB€1.73 / GPU-hrEUOn request
NVIDIA RTX 5090 ×4 — 32 GB each€0.54 / GPU-hrEUOn request
NVIDIA RTX 3090 ×4 — 24 GB each€0.17 / GPU-hrEUOn request
Model deployment — validated open models or your own weightsRequest quoteEUOn request

Deployment and operation included — machine, CUDA, runtime, endpoint. An engineer confirms the configuration and monthly price before anything is billed. Request deployment.

GPU VM

The GPU, your stack

GPUMemoryPriceRegionStatus
DGX Spark (GB10)128 GB unified€495 / montheu-se-1Available
RTX 6000 Pro96 GB€1.73 / GPU-hrEUOn request
RTX 5090 ×432 GB each€0.54 / GPU-hrEUOn request
RTX 3090 ×424 GB each€0.17 / GPU-hrEUOn request
H100 / H200 — against registered demand80 / 141 GBEUOn request

SSH access, dedicated machine, base environment included. Hardware without a price carries none until it lands. Deploy GPU.

The pricing rule

80% of the lowest proper-supplier price

We price inference aggressively low: our target is 80% of the lowest current price from a proper inference supplier for the same model, tracked continuously. Found better pricing at a proper inference supplier? Tell us and we’ll look into lowering ours.

In practice: we maintain a price index of suppliers serving the same model and set our target at 80% of the lowest entry. A proper inference supplier is a provider operating and standing behind managed inference for the model — decentralized compute marketplaces don't count. The current Qwen3.8 27B prices (€0.29 / €1.77 per 1M tokens) were set against the index of 2026-08-26.

Why so cheap

Why so cheap, honestly

Three reasons, none of them a trick:

We own the hardwareInference runs on NVIDIA DGX Spark systems AxForge owns and operates in Sweden — no cloud markup between you and the silicon.
Efficient models, right-sized machinesWe serve efficient open-weight models on machines sized for exactly that job, so the hardware is actually utilized instead of idling expensively.
Real cost plus a thin marginThe price reflects what serving actually costs us, plus a thin margin. No VC-subsidized loss-leader pricing that has to snap back later.

What the low price does not buy us: your data. Prompts and completions are processed in memory, never retained, never used to train — see the privacy policy.

FAQ

Pricing — common questions

Is there a free tier?

Not a public one today. API keys are allocated from the waiting list, and pay-as-you-go starts at the listed per-token prices — no subscription, no minimum.

What does "launch pricing" mean?

These are the opening prices for the shared API, set by our pricing rule — 80% of the lowest current price from a proper inference supplier for the same model at the time we set them (index of 2026-08-26). It is not a teaser with a preset expiry: future prices follow the same rule, tracked against the market.

How does the 80% pricing rule work?

We track a price index of proper inference suppliers serving the same model and target 80% of the lowest current entry, tracked continuously. If you find better pricing at a proper inference supplier, tell us and we’ll look into lowering ours.

Do prices change?

They can, in both directions, because the target tracks the market. What never changes with the price: EU residency, zero prompt retention, and no training on customer data.

Is billing per token?

Chat is billed per token with separate input and output rates; embeddings per input token; image generation per image. Every API response reports usage — including in streaming mode — so you can reconcile your bill exactly. List prices exclude VAT.

What about volume?

There are no published volume tiers. For sustained high volume, a dedicated system — a monthly DGX Spark rental serving only your traffic — is usually the better shape. Talk to an engineer and we’ll quote it.

Related

Nearby on AxForge

Transparent AI pricing, from Europe, tracked against the market.

Get an API key Talk to an engineer
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms