Model catalogue · api.axforge.ai/v1

Open models served from Europe

Open-weight models behind one OpenAI-compatible API, running on hardware AxForge owns in the EU. Seven models are served now on the shared API; six more deploy on dedicated systems on request.

eu-se-1 · Stockholm eu-es-1 · Málaga More EU regions — in deployment
Get an API key Talk to an engineer Keys are allocated from the waiting list.

Serverless Models

Available through Serverless Models

ModelTypeAvailabilityBest forPrice
Qwen3.8 27BLLMAvailableAssistants, tool calling, RAG — the daily driver€0.29 in · €1.77 out / 1MGet API key
Qwen3 EmbeddingEmbeddingAvailableSemantic search and retrieval€0.015 / 1MGet API key
ERNIE Image TurboImage generationAvailablePosters, banners and UI shots with readable text€0.02 / imageGet API key
FLUX.2 Klein 4BImage editingAvailablePhotoshop-style edits from one sentence, ~8 s€0.02 / imageGet API key
MiniMax Music 3Music generationAvailableFull songs and soundtracks from a prompt€0.09 / trackGet API key
WhisperSpeech-to-textAvailableTranscription and voice input€0.005 / minuteGet API key
PiperText-to-speechAvailableFast voice output€2.95 / 1M charactersGet API key

Managed GPU

Available for Managed GPU deployment

Validated on AxForge hardware and brought hot on a dedicated system for you — not on the shared API today, so no shared-API price is listed.

ModelTypeAvailabilityBest forPrice
Qwen-ImageImage gen + editOn requestGeneration and editing on one dedicated machineOn requestRequest
Qwen3.6 35B A3BLLM (MoE)On requestHigh-throughput chat on a dedicated SparkOn requestRequest
Qwen3 30B A3BLLM (MoE)On requestEfficient volume workloadsOn requestRequest
Gemma-4 26BLLM (vision)On requestLong documents — 262,144-token context — plus visionOn requestRequest
Mistral Small 3.2LLMOn requestCompact European-family workhorseOn requestRequest
SDXLImage generationOn requestHigh-volume light image generationOn requestRequest
  • Pricing rule: we target 80% of the lowest current price from a proper inference supplier for the same model, tracked continuously. Found better pricing at a proper inference supplier? Tell us — we’ll look into lowering ours.
  • Qwen3.8 27B pricing is launch pricing. List prices are in EUR and exclude VAT.
  • Every served model answers on one OpenAI-compatible API — api.axforge.ai/v1 — with EU region pinning on the key.

Two products

Serverless Models or Managed GPU

Serverless Models · Available

The shared API. One key, OpenAI-compatible endpoints for chat, embeddings, images and audio. Streaming, tool calls and usage reporting. Requests are pinned to an EU region on the key and echoed on the response.

Managed GPU · On request

A model validated on AxForge hardware, brought hot on a dedicated NVIDIA DGX Spark as a monthly rental. An OpenAI-compatible endpoint on your own machine, serving only your traffic. See the DGX Spark rental page.

Data & privacy

Zero prompt retention

Prompts and completions are processed in memory in Sweden — not written to disk, not logged, not retained, and never used to train anything. We keep only request metadata (token counts, timestamps, status) for billing and operations. The full policy is at axforge.ai/privacy.

FAQ

The model catalogue — common questions

Which models are on the shared API today?

Seven: Qwen3.8 27B (chat), Qwen3 Embedding, ERNIE Image Turbo (image generation), FLUX.2 Klein 4B (image editing), MiniMax Music 3 (music generation), Whisper (speech-to-text) and Piper (text-to-speech) — all behind one OpenAI-compatible API at api.axforge.ai/v1.

What does "deployable on request" mean?

The model is validated on AxForge hardware and can be brought hot on a dedicated NVIDIA DGX Spark as a monthly rental, with an OpenAI-compatible endpoint serving only your traffic. It is not on the shared API today. Talk to an engineer to start.

Where do the models run?

In Stockholm, Sweden (region eu-se-1), on hardware AxForge owns and operates. Spain (eu-es-1) is live for the platform; more EU regions are in deployment. TLS terminates in the EU.

Are prompts retained or used for training?

No. Prompts and completions are processed in memory in Sweden — not written to disk, not logged, not retained, never used to train. Only request metadata is kept for billing and operations — see the privacy policy.

What do the model APIs cost?

Qwen3.8 27B is €0.29 per 1M input tokens and €1.77 per 1M output tokens (launch pricing). Embeddings are €0.015 per 1M tokens, images €0.02 each, music €0.09 per track, speech-to-text €0.005 per minute and text-to-speech €2.95 per 1M characters. Dedicated deployments are priced on request — talk to an engineer.

Related

Nearby on AxForge

One API. EU-owned hardware. Zero retention.

Get an API key Talk to an engineer
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms