Model reference · open weights

Qwen3-Next-Thinking

Available as managed deployment LLMs Qwen Text gen · MoE 2 variants 45k dl/mo

Qwen3-Next-Thinking is an open-weight language model from Qwen. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byQwen
TypeLanguage models
TaskText gen · MoE
Parameters (lead)81.3B
Context256k tokens
Runs withtransformers
Released2025-09-09
Popularity45k downloads / month
LicenceOpen weights

About

What Qwen3-Next-Thinking is

Over the past few months, we have observed increasingly clear trends toward scaling both total parameters and context lengths in the pursuit of more powerful and agentic artificial intelligence (AI). We are excited to share our latest advancements in addressing these demands, centered on improving scaling efficiency through innovative model architecture. We call this next-generation foundation models Qwen3-Next.

Read the full model card

Highlights

Qwen3-Next-80B-A3B is the first installment in the Qwen3-Next series and features the following key enchancements:

  • Hybrid Attention: Replaces standard attention with the combination of Gated DeltaNet and Gated Attention, enabling efficient context modeling for ultra-long context length.
  • High-Sparsity Mixture-of-Experts (MoE): Achieves an extreme low activation ratio in MoE layers, drastically reducing FLOPs per token while preserving model capacity.
  • Stability Optimizations: Includes techniques such as zero-centered and weight-decayed layernorm, and other stabilizing enhancements for robust pre-training and post-training.
  • Multi-Token Prediction (MTP): Boosts pretraining model performance and accelerates inference.

We are seeing strong performance in terms of both parameter efficiency and inference speed for Qwen3-Next-80B-A3B:

  • Qwen3-Next-80B-A3B-Base outperforms Qwen3-32B-Base on downstream tasks with 10% of the total training cost and with 10 times inference throughput for context over 32K tokens.
  • Leveraging GSPO, we have addressed the stability and efficiency challenges posed by the hybrid attention mechanism combined with a high-sparsity MoE architecture in RL training. Qwen3-Next-80B-A3B-Thinking demonstrates outstanding performance on complex reasoning tasks, not only surpassing Qwen3-30B-A3B-Thinking-2507 and Qwen3-32B-Thinking, but also outperforming the proprietary model Gemini-2.5-Flash-Thinking across multiple benchmarks.

For more details, please refer to our blog post Qwen3-Next.

Model Overview

[!Note] Qwen3-Next-80B-A3B-Thinking supports only thinking mode. To enforce model thinking, the default chat template automatically includes . Therefore, it is normal for the model's output to contain only without an explicit opening `` tag.

[!Note] Qwen3-Next-80B-A3B-Thinking may generate thinking content longer than its predecessor. We strongly recommend its use in highly complex reasoning tasks.

Qwen3-Next-80B-A3B-Thinking has the following features:

  • Type: Causal Language Models
  • Training Stage: Pretraining (15T tokens) & Post-training
  • Number of Parameters: 80B in total and 3B activated
  • Number of Paramaters (Non-Embedding): 79B
  • Hidden Dimension: 2048
  • Number of Layers: 48
    • Hybrid Layout: 12 * (3 * (Gated DeltaNet -> MoE) -> 1 * (Gated Attention -> MoE))
  • Gated Attention:
    • Number of Attention Heads: 16 for Q and 2 for KV
    • Head Dimension: 256
    • Rotary Position Embedding Dimension: 64
  • Gated DeltaNet:
    • Number of Linear Attention Heads: 32 for V and 16 for QK
    • Head Dimension: 128
  • Mixture of Experts:
    • Number of Experts: 512
    • Number of Activated Experts: 10
    • Number of Shared Experts: 1
    • Expert Intermediate Dimension: 512
  • Context Length: 262,144 natively and extensible up to 1,010,000 tokens

Performance

Qwen3-30B-A3B-Thinking-2507Qwen3-32B ThinkingQwen3-235B-A22B-Thinking-2507Gemini-2.5-Flash ThinkingQwen3-Next-80B-A3B-Thinking
Knowledge
MMLU-Pro80.979.184.481.982.7
MMLU-Redux91.490.993.892.192.5
GPQA73.468.481.182.877.2
SuperGPQA56.854.164.957.860.8
Reasoning
AIME2585.072.992.372.087.8
HMMT2571.451.583.964.273.9
LiveBench 24112576.874.978.474.376.6
Coding
LiveCodeBench v6 (25.02-25.05)66.060.674.161.268.7
CFEval20441986213419952071
OJBench25.124.132.523.529.7
Alignment
IFEval88.985.087.889.888.9
Arena-Hard v2*56.048.479.756.762.3
WritingBench85.079.088.383.984.6
Agent
BFCL-v372.470.371.968.672.0
TAU1-Retail67.852.867.865.269.6
TAU1-Airline48.029.046.054.049.0
TAU2-Retail58.849.771.966.767.8
TAU2-Airline58.045.558.052.060.5
TAU2-Telecom26.327.245.631.643.9
Multilingualism
MultiIF76.473.080.674.477.8
MMLU-ProX76.474.681.080.278.7
INCLUDE74.473.781.083.978.9
PolyMATH52.647.460.149.856.3

*: For reproducibility, we report the win rates evaluated by GPT-4.1.

Quickstart

The code for Qwen3-Next has been merged into the main branch of Hugging Face transformers.

pip install git+https://github.com/huggingface/transformers.git@main

With earlier versions, you will encounter the following error:

KeyError: 'qwen3_next'

The following contains a code snippet illustrating how to use the model generate content based on given inputs.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen3-Next-80B-A3B-Thinking"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    dtype="auto",
    device_map="auto"
)

# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
    {"role": "user", "content": prompt},
]
text = t

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys qwen3-next-thinking for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (qwen3-next-thinking below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-next-thinking","messages":[{"role":"user","content":"Hello"}]}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms