GPU rental · available now
A dedicated GB10 Grace-Blackwell system with 128 GB unified memory, on monthly rental from hardware AxForge owns and operates in the EU. Your model, your traffic, our machine.
Specifications
| System | NVIDIA DGX Spark — dedicated, single-tenant |
|---|---|
| Superchip | NVIDIA GB10 Grace-Blackwell — see the GB10 page |
| Memory | 128 GB unified, shared CPU/GPU |
| CPU architecture | ARM64 |
| Availability | Available — 8 units in Sweden (eu-se-1); 3 more on the way — days |
| Rental term | Monthly |
| Pricing | €495 / month (launch pricing) |
| Region | eu-se-1 · Stockholm, Sweden |
Performance
| Workload | Result | Condition |
|---|---|---|
| Qwen3.8 27B decode | 15.9 tokens/s | single stream, multi-token-prediction speculative decoding on (5.8 tokens/s without) |
| Qwen3.6 35B A3B decode | 30.7 tokens/s | single stream; MoE, 3B active parameters |
| ERNIE Image Turbo 1024×1024 | ~32 s / image | 8-step turbo fp8, current generation model |
| Qwen-Image 1024×1024 | ~26 s / image | 8-step lightning (now serving edits) |
Measured on our production DGX Spark node, single-stream, 2026-08. Serving config for Qwen3.8 27B: 65,536 context, up to 4 concurrent sequences. We publish only numbers we measured ourselves.
Fit
| Use case | Why it fits |
|---|---|
| Dedicated inference, ~7B–35B models | 128 GB unified memory holds model weights and KV cache in one pool — the measured numbers above are exactly this workload. |
| Private inference | A single-tenant machine in an EU region. Your model, your traffic, our hardware — prompts never persisted. |
| Dev / staging nodes | A named machine you keep for the month — a stable target for integration, load testing and pre-production serving. |
Searching for DGX Spark cloud or DGX Spark hosting? Same machine: you rent a named physical system, we host and operate it, you reach it over the network.
Data & privacy
Prompts never persisted. Requests to your dedicated DGX Spark are processed in memory in Sweden — not written to disk, not logged, not retained, never used to train anything. We keep only request metadata (token counts, timestamps, status) for billing and operations. Full policy at axforge.ai/privacy.
FAQ
Yes. 8 units are available now in Sweden (eu-se-1) on monthly rental, and 3 more are arriving in days. Talk to an engineer to claim one.
€495 per month for a dedicated machine (launch pricing). We watch the market and price under it: that is 90% of the lowest listed dedicated DGX Spark rental we found on 2026-08-26. An engineer scopes the configuration with you.
Dedicated inference of roughly 7B to 35B models. On our production nodes we measured Qwen3.8 27B at 15.9 tokens/s single-stream with speculative decoding and Qwen3.6 35B A3B at 30.7 tokens/s single-stream (2026-08). See the GB10 page for the full table.
On AxForge you rent a named physical machine, hosted and operated by us in an EU region, reachable over the network like any cloud endpoint — but it is your dedicated system, not a shared cloud instance.
The GB10 platform is ARM64. Our own serving stack runs on it in production — the published numbers were measured there. x86-only binaries need ARM64 builds; an engineer can review your stack before you commit.
No. Your model, your traffic, our hardware — prompts never persisted. Only request metadata (token counts, timestamps, status) is kept for billing and operations — see the privacy policy.