Dedicated model · Available as managed deployment
Google's Gemma 4 31B — the leading dense open model built to run strongly on a single GPU, with text and image input, under Apache-2.0. Available as a managed deployment: validated on AxForge hardware and deployed on a DGX Spark for your traffic only — an OpenAI-compatible endpoint on hardware only you use, hosted in the EU.
Why Gemma 4 31B
| Dense and single-GPU | 31B dense parameters designed for one accelerator: on a dedicated DGX Spark it is the strongest single-machine option in its class. |
|---|---|
| Text + vision | Accepts images alongside text — screenshots, scans and photos through the same endpoint. |
| Apache-2.0 | Permissive weights: deploy, fine-tune and ship commercially without a licence conversation. |
Specifications
| Model | Gemma 4 31B — google |
|---|---|
| Modalities | Text + vision — accepts image input |
| Sizes | 12.0B, 25.8B, 31.3B, 32.7B |
| Released | 2026-03 |
| Popularity | 8.5M downloads / month on HuggingFace |
| Licence | Open weights — apache-2.0; commercial use permitted |
| Availability | Available as managed deployment — operated by AxForge on hardware reserved for you |
| Hardware | NVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge |
| Region | Málaga, Spain (eu-es-1) |
| Pricing | Hardware from €0.85/hour — rental term: Hour, week, month or year. Managed service quoted per deployment — Request deployment and we confirm both in writing. |
We publish only numbers we measure ourselves, and we haven't benchmarked this model on our nodes yet — for quality benchmarks, see the model card in the catalogue and the official card on HuggingFace.
Deployment
| System | A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory), reserved for you |
|---|---|
| Commercial model | Hour, week, month or year — hardware from €0.85/hour, managed service quoted per deployment |
| Endpoint | OpenAI-compatible /v1 on your own machine — only your traffic |
| Location | Hosted in the EU — Málaga, Spain (eu-es-1) |
| Data handling | Zero prompt retention, same policy as the serverless API |
The model is validated on AxForge hardware and deployed on your system. After deployment, your endpoint speaks the OpenAI API with the model name you receive with the deployment.
Data & privacy
Prompts and completions are processed in memory in the EU — not written to disk, not logged, not retained, and never used to train anything. We keep only request metadata (token counts, timestamps, status) for billing and operations. The full policy is at axforge.ai/privacy.
FAQ
Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.
31B is the dense model — best quality per machine. 26B-A4B is the mixture-of-experts sibling AxForge also offers, faster per token; both run on a dedicated DGX Spark.
Yes — natively, with room for its full context; AxForge validates the build with the deployment.
We publish only numbers we measure ourselves, and we haven't benchmarked this model on our nodes yet. For quality benchmarks, see the official model card.
A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) rented by the hour, week, month or year, running Gemma 4 31B behind an OpenAI-compatible endpoint on your own machine, hosted in the EU with zero prompt retention.
Hardware from €0.85/hour, by the hour, week, month or year; the managed service is quoted per deployment. Request deployment and we confirm both in writing.
Explore