Dedicated model · deployable on request

Qwen3 30B A3B — dedicated hosting in Europe

A mixture-of-experts Qwen3, validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. Not on the shared API today.

On request eu-se-1 · Stockholm 65,536 ctx · MoE
Talk to an engineer Qwen3.8 27B — served now Monthly rental. Pricing on request.

Specifications

The model

ModelQwen3 30B A3B — mixture-of-experts, Qwen3 family
Served model nameqwen3-30b-a3b
ArchitectureMoE
Context window65,536 tokens
AvailabilityDeployable on request — not on the shared API today
HardwareNVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge
Regioneu-se-1 · Stockholm, Sweden
PricingOn request — talk to an engineer

We publish only numbers we measure ourselves, and we haven't benchmarked this model on our nodes yet — for quality benchmarks, see the official model card.

Deployment

What "deployable on request" means

SystemA dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory), reserved for you
Commercial modelMonthly rental
EndpointOpenAI-compatible /v1 on your own machine — only your traffic
LocationHosted in the EU (eu-se-1 · Stockholm, Sweden)
Data handlingZero prompt retention, same policy as the shared API

The model is validated on AxForge hardware and brought hot on your system. After deployment — and only then — your endpoint speaks the OpenAI API with the model name qwen3-30b-a3b.

Data & privacy

Zero prompt retention

Prompts and completions are processed in memory in Sweden — not written to disk, not logged, not retained, and never used to train anything. We keep only request metadata (token counts, timestamps, status) for billing and operations. The full policy is at axforge.ai/privacy.

FAQ

Qwen3 30B A3B hosting — common questions

Is Qwen3 30B A3B on the shared AxForge API?

No. It is not on the shared API today. It is deployable on request: validated on AxForge hardware and brought hot on a dedicated DGX Spark for your traffic only. The shared API serves Qwen3.8 27B.

What kind of model is Qwen3 30B A3B?

A mixture-of-experts LLM from the Qwen3 family with a 65,536-token context window.

How fast is Qwen3 30B A3B on your hardware?

We publish only numbers we measure ourselves, and we haven't benchmarked this model on our nodes yet. For quality benchmarks, see the official model card — the newer Qwen3.6 35B A3B has a measured figure on its page.

What does a dedicated deployment look like?

A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) rented monthly, running Qwen3 30B A3B behind an OpenAI-compatible endpoint on your own machine, hosted in the EU with zero prompt retention.

What does Qwen3 30B A3B hosting in Europe cost?

Pricing on request — talk to an engineer and we will quote the monthly rental.

Ready to build on EU inference?

Request deployment Talk to an engineer

Explore

All models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms