Comparison · different shapes of GPU cloud
RunPod is a large self-serve GPU cloud built for instant scale; AxForge is a small EU operator renting named machines by the month, with a zero-retention hosted model API.
At a glance
| Dimension | AxForge | RunPod |
|---|---|---|
| What it is | A small EU infrastructure operator: dedicated machines plus a hosted open-model API, on hardware AxForge owns. | A large self-serve GPU cloud: rent GPU capacity on demand, at scale. |
| How you get compute | Request a system and talk to an engineer; API keys are allocated from a waiting list. Engineer-led onboarding. | Sign up and launch in minutes, fully self-serve. |
| GPU menu | Short and named: NVIDIA DGX Spark (GB10, 128 GB unified memory) available now; RTX 6000 Pro, RTX 5090 and RTX 3090 tiers on the way. | A very large menu of GPU types across on-demand pods and serverless workers. |
| Billing model | Monthly rental for dedicated machines; per-token billing on the hosted API. | Per-second, on-demand billing — spin up and down as you need. |
| Hosted inference | A managed OpenAI-compatible model API — Qwen3.8 27B served now, plus embeddings, image and speech — with zero prompt retention. | Bring your own containers and models; serverless workers scale them for you. |
| EU data residency guarantee | EU-only by construction: eu-se-1 Stockholm and eu-es-1 Málaga live, region pinned per key and echoed on the response. | Datacenters in many regions worldwide; you choose where a workload runs. |
| Hardware transparency | Named, owned machines with published serving details and performance we measured ourselves. | You pick a GPU SKU from the menu; the platform operates the fleet. |
Honest routing
A large self-serve GPU cloud wins on these requirements:
| Compute in minutes | You need a GPU now, self-serve, with no conversation and no waiting list. |
|---|---|
| Burst capacity | Spiky training or batch workloads where per-second billing and instant scale-up beat a monthly commitment. |
| A specific GPU SKU | You want a particular card from a very large menu of GPU types. |
| Serverless workers | You want your own containers scaled up and down automatically, including to zero. |
Honest routing
| EU sovereignty requirements | Processing must stay in the EU and you must be able to prove it — named machine, named region, pinned per key, echoed on every response. |
|---|---|
| Predictable dedicated capacity | One named machine, your traffic only, a flat monthly rental — no bidding, no preemption, no noisy neighbours. |
| A hosted model API | You want open-model inference as a managed API with zero prompt retention, instead of operating your own serving stack. |
| Auditability | You need to tell an auditor exactly what hardware ran the workload and how it was configured — and get answers from the engineers who run it. |
The substance
| Regions | eu-se-1 · Stockholm and eu-es-1 · Málaga, both live. More EU regions in deployment. |
|---|---|
| Data handling | Zero prompt retention — in-memory processing in Sweden, no disk, no logs, no training. Only billing metadata is kept. Full policy: axforge.ai/privacy. |
| API | OpenAI-compatible at https://api.axforge.ai/v1 — TLS 1.3, terminated in the EU. |
| Pricing | Qwen3.8 27B: €0.29 / 1M input · €1.77 / 1M output (launch pricing). GPU rentals: €495 / month for a dedicated DGX Spark — full price list published. |
| Measured speed | 15.9 tokens/s single-stream decode for Qwen3.8 27B with speculative decoding. |
| Dedicated systems | NVIDIA DGX Spark (GB10, 128 GB unified memory), monthly rental, available now in Sweden. RTX tiers on the way. |
Speed measured on our production DGX Spark node, single-stream, 2026-08. We publish only numbers we measured ourselves.
FAQ
For EU sovereignty requirements and predictable dedicated capacity, yes: AxForge rents named machines monthly from EU regions and runs a zero-retention model API. For instant self-serve scale, burst workloads and a huge GPU menu, a large GPU cloud like RunPod is the better fit.
Yes. NVIDIA DGX Spark systems (GB10, 128 GB
unified memory) are available now as monthly rentals in Sweden
(eu-se-1). RTX 6000 Pro, RTX 5090 and RTX 3090 tiers are on the
way. Onboarding is engineer-led — you talk to the people running the hardware.
No. Dedicated machines are monthly rentals, and the hosted model API is billed per token. If you need per-second billing and scale-to-zero, a self-serve GPU cloud is the right tool.
Qwen3.8 27B is €0.29 per million input tokens and €1.77 per million output tokens (launch pricing). GPU rentals carry published launch pricing too — a dedicated DGX Spark is €495 per month, and the full price list is public. Talk to an engineer to scope a system.
In the EU only: eu-se-1 (Stockholm) and eu-es-1
(Málaga) are live, with more EU regions in deployment. The region is pinned per
API key and echoed on every response, and prompts are processed with
zero retention.
Related