Dedicated model · Available as managed deployment

Gemma 4 31B API in Europe — managed deployment

Google's Gemma 4 31B — the leading dense open model built to run strongly on a single GPU, with text and image input, under Apache-2.0. Available as a managed deployment: validated on AxForge hardware and deployed on a DGX Spark for your traffic only — an OpenAI-compatible endpoint on hardware only you use, hosted in the EU.

Available as managed deployment eu-es-1 · Málaga Text + vision
Request deployment Qwen3.8 27B — on the API now Hour, week, month or year. Hardware from €0.85/hour; managed service quoted per deployment.

Why Gemma 4 31B

What you get

Dense and single-GPU31B dense parameters designed for one accelerator: on a dedicated DGX Spark it is the strongest single-machine option in its class.
Text + visionAccepts images alongside text — screenshots, scans and photos through the same endpoint.
Apache-2.0Permissive weights: deploy, fine-tune and ship commercially without a licence conversation.

Specifications

The model

ModelGemma 4 31B — google
ModalitiesText + vision — accepts image input
Sizes12.0B, 25.8B, 31.3B, 32.7B
Released2026-03
Popularity8.5M downloads / month on HuggingFace
LicenceOpen weights — apache-2.0; commercial use permitted
AvailabilityAvailable as managed deployment — operated by AxForge on hardware reserved for you
HardwareNVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge
RegionMálaga, Spain (eu-es-1)
PricingHardware from €0.85/hour — rental term: Hour, week, month or year. Managed service quoted per deployment — Request deployment and we confirm both in writing.

We publish only numbers we measure ourselves, and we haven't benchmarked this model on our nodes yet — for quality benchmarks, see the model card in the catalogue and the official card on HuggingFace.

Deployment

What "managed deployment" means

SystemA dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory), reserved for you
Commercial modelHour, week, month or year — hardware from €0.85/hour, managed service quoted per deployment
EndpointOpenAI-compatible /v1 on your own machine — only your traffic
LocationHosted in the EU — Málaga, Spain (eu-es-1)
Data handlingZero prompt retention, same policy as the serverless API

The model is validated on AxForge hardware and deployed on your system. After deployment, your endpoint speaks the OpenAI API with the model name you receive with the deployment.

Data & privacy

Zero prompt retention

Prompts and completions are processed in memory in the EU — not written to disk, not logged, not retained, and never used to train anything. We keep only request metadata (token counts, timestamps, status) for billing and operations. The full policy is at axforge.ai/privacy.

FAQ

Gemma 4 31B hosting — common questions

Is Gemma 4 31B on the AxForge serverless API?

Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.

Gemma 4 31B or Gemma 4 26B?

31B is the dense model — best quality per machine. 26B-A4B is the mixture-of-experts sibling AxForge also offers, faster per token; both run on a dedicated DGX Spark.

Does Gemma 4 31B fit on a DGX Spark?

Yes — natively, with room for its full context; AxForge validates the build with the deployment.

How fast is Gemma 4 31B on your hardware?

We publish only numbers we measure ourselves, and we haven't benchmarked this model on our nodes yet. For quality benchmarks, see the official model card.

What does a dedicated Gemma 4 31B deployment look like?

A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) rented by the hour, week, month or year, running Gemma 4 31B behind an OpenAI-compatible endpoint on your own machine, hosted in the EU with zero prompt retention.

What does EU Gemma 4 31B hosting cost?

Hardware from €0.85/hour, by the hour, week, month or year; the managed service is quoted per deployment. Request deployment and we confirm both in writing.

Ready to build on EU inference?

Request deployment Talk to an engineer

Explore

All models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms