Dedicated model · deployable on request
A vision-capable open model with a 262,144-token context window, validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. Not on the shared API today.
Specifications
| Model | Gemma-4 26B — open model, Gemma family |
|---|---|
| Served model name | gemma |
| Context window | 262,144 tokens |
| Modalities | Text + vision — accepts image input |
| Availability | Deployable on request — not on the shared API today |
| Hardware | NVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge |
| Region | eu-se-1 · Stockholm, Sweden |
| Pricing | On request — talk to an engineer |
We publish only numbers we measure ourselves, and we haven't benchmarked this model on our nodes yet — for quality benchmarks, see the official model card.
Deployment
| System | A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory), reserved for you |
|---|---|
| Commercial model | Monthly rental |
| Endpoint | OpenAI-compatible /v1 on your own machine — only your traffic |
| Location | Hosted in the EU (eu-se-1 · Stockholm, Sweden) |
| Data handling | Zero prompt retention, same policy as the shared API |
The model is validated on AxForge hardware and brought hot on your system. After deployment — and only then — your endpoint speaks the OpenAI API with the model name gemma.
Data & privacy
Prompts and completions are processed in memory in Sweden — not written to disk, not logged, not retained, and never used to train anything. We keep only request metadata (token counts, timestamps, status) for billing and operations. The full policy is at axforge.ai/privacy.
FAQ
No. It is not on the shared API today. It is deployable on request: validated on AxForge hardware and brought hot on a dedicated DGX Spark for your traffic only. The shared API serves Qwen3.8 27B.
262,144 tokens — the largest context window in the AxForge catalogue.
Yes, the model is vision-capable: it accepts image input alongside text.
We publish only numbers we measure ourselves, and we haven't benchmarked this model on our nodes yet. For quality benchmarks, see the official model card.
A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) rented monthly, running Gemma-4 26B behind an OpenAI-compatible endpoint on your own machine, hosted in the EU with zero prompt retention.
Pricing on request — talk to an engineer and we will quote the monthly rental.