Model reference · open weights

Qwen3.6

Qwen3.6 is an open-weight language model from nvidia, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

LLMs nvidia 2 variants 11.2M downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What Qwen3.6 is

Model Overview Description: The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. <br Third-Party Community Consideration This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.6-35B-A3B) Model Card from Alibaba. References NVIDIA Model Optimizer: https://github.com/NVIDIA/Model-Optimizer License/Terms of Use: GOVERNING DOWNLOAD TERMS: Use of the model is governed by the Apache license 2.0. Deployment Geography: Global <br Use Case: <br Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG systems, and other AI-powered applications. <br Release Date: <br Hugging Face on 05/28/2026 via https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4 <br Model Architecture: Architecture Type: Transformers <br Network Architecture: Mixture-of-Experts (MoE) with Hybrid Attention <br Number of Model Parameters: 35B in total and 3B activated <br Input: Input Type(s): Text, Image, Video <br Input Format(s): String, Red, Green, Blue (RGB), Video (MP4/WebM) <br Input Parameters: One-Dimensional (1D), Two-Dimensional (2D), Three-Dimensional (3D) <br Other Properties Related to Input: Context length up to 262K <br Output: Output Type(s): Text <br Output Format: String <br Output Parameters: One-Dimensional(1D): Sequences <br Other Properties Related to Output: None <br Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions. <br Software Integration: Supported Runtime Engine(s): <br vLLM <br Supported Hardware Microarchitecture Compatibility: <br NVIDIA Hopper, NVIDIA Blackwell <br Preferr

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makernvidia
TypeLanguage models
Parameters (lead)18.7B
Variants2
Runs withModel Optimizer
Based onQwen/Qwen3.6-35B-A3B
Released2026-05-27
Popularity11.2M downloads / month
Likes576
LicenceOpen weights

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
Qwen3.6-35B-A3B-NVFP418.7BNVFP4Weights ↗
Qwen3.6-27B-NVFP418.2BNVFP4Weights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys nvidia-qwen3-6 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (nvidia-qwen3-6 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia-qwen3-6","messages":[{"role":"user","content":"Hello"}]}'

Details

Languages, data & research

Tags

Model Optimizer safetensors qwen3_5_moe nvidia ModelOpt Qwen3.6 quantized FP4 fp4 text-generation conversational 8-bit modelopt deploy:azure

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want Qwen3.6 on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More language models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms