Model catalogue · api.axforge.ai/v1
Open-weight models behind one OpenAI-compatible API, running on hardware AxForge owns in the EU. Seven models are served now on the shared API; six more deploy on dedicated systems on request.
Serverless Models
| Model | Type | Availability | Best for | Price | |
|---|---|---|---|---|---|
| Qwen3.8 27B | LLM | Available | Assistants, tool calling, RAG — the daily driver | €0.29 in · €1.77 out / 1M | Get API key |
| Qwen3 Embedding | Embedding | Available | Semantic search and retrieval | €0.015 / 1M | Get API key |
| ERNIE Image Turbo | Image generation | Available | Posters, banners and UI shots with readable text | €0.02 / image | Get API key |
| FLUX.2 Klein 4B | Image editing | Available | Photoshop-style edits from one sentence, ~8 s | €0.02 / image | Get API key |
| MiniMax Music 3 | Music generation | Available | Full songs and soundtracks from a prompt | €0.09 / track | Get API key |
| Whisper | Speech-to-text | Available | Transcription and voice input | €0.005 / minute | Get API key |
| Piper | Text-to-speech | Available | Fast voice output | €2.95 / 1M characters | Get API key |
Managed GPU
Validated on AxForge hardware and brought hot on a dedicated system for you — not on the shared API today, so no shared-API price is listed.
| Model | Type | Availability | Best for | Price | |
|---|---|---|---|---|---|
| Qwen-Image | Image gen + edit | On request | Generation and editing on one dedicated machine | On request | Request |
| Qwen3.6 35B A3B | LLM (MoE) | On request | High-throughput chat on a dedicated Spark | On request | Request |
| Qwen3 30B A3B | LLM (MoE) | On request | Efficient volume workloads | On request | Request |
| Gemma-4 26B | LLM (vision) | On request | Long documents — 262,144-token context — plus vision | On request | Request |
| Mistral Small 3.2 | LLM | On request | Compact European-family workhorse | On request | Request |
| SDXL | Image generation | On request | High-volume light image generation | On request | Request |
Two products
Serverless Models · Available
The shared API. One key, OpenAI-compatible endpoints for chat, embeddings, images and audio. Streaming, tool calls and usage reporting. Requests are pinned to an EU region on the key and echoed on the response.
Managed GPU · On request
A model validated on AxForge hardware, brought hot on a dedicated NVIDIA DGX Spark as a monthly rental. An OpenAI-compatible endpoint on your own machine, serving only your traffic. See the DGX Spark rental page.
Data & privacy
Prompts and completions are processed in memory in Sweden — not written to disk, not logged, not retained, and never used to train anything. We keep only request metadata (token counts, timestamps, status) for billing and operations. The full policy is at axforge.ai/privacy.
FAQ
Seven: Qwen3.8 27B (chat),
Qwen3 Embedding,
ERNIE Image Turbo (image generation),
FLUX.2 Klein 4B (image editing),
MiniMax Music 3 (music generation),
Whisper (speech-to-text) and
Piper (text-to-speech) — all
behind one OpenAI-compatible API at api.axforge.ai/v1.
The model is validated on AxForge hardware and can be brought hot on a dedicated NVIDIA DGX Spark as a monthly rental, with an OpenAI-compatible endpoint serving only your traffic. It is not on the shared API today. Talk to an engineer to start.
In Stockholm, Sweden (region eu-se-1), on hardware AxForge
owns and operates. Spain (eu-es-1) is live for the platform;
more EU regions are in deployment. TLS terminates in the EU.
No. Prompts and completions are processed in memory in Sweden — not written to disk, not logged, not retained, never used to train. Only request metadata is kept for billing and operations — see the privacy policy.
Qwen3.8 27B is €0.29 per 1M input tokens and €1.77 per 1M output tokens (launch pricing). Embeddings are €0.015 per 1M tokens, images €0.02 each, music €0.09 per track, speech-to-text €0.005 per minute and text-to-speech €2.95 per 1M characters. Dedicated deployments are priced on request — talk to an engineer.