Model reference · open weights

Qwen2.5

Qwen2.5 is an open-weight language model from Qwen, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Licence fee required LLMs Qwen 8 variants 10.8M downloads/mo
Request a licence + hosting quote All served models Not on the shared API today — deployed on request.

About

What Qwen2.5 is

Qwen2.5-7B-Instruct Introduction Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains. - Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and condition-setting for chatbots. - Long-context Support up to 128K tokens and can generate up to 8K tokens. - Multilingual support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more. This repo contains the instruction-tuned 7B Qwen2.5 model, which has the following features: - Type: Causal Language Models - Training Stage: Pretraining & Post-training - Architecture: transformers with RoPE, SwiGLU, RMSNorm, and Attention QKV bias - Number of Parameters: 7.61B - Number of Paramaters (Non-Embedding): 6.53B - Number of Layers: 28 - Number of Attention Heads (GQA): 28 for Q and 4 for KV - Context Length: Full 131,072 tokens and generation 8192 tokens - Please refer to this section for detailed instructions on how to deploy Qwen2.5 for handling long texts. For more details, please refer to our blog, GitHub, and Documentation. Requirements The code of Qwen2.5 has been in the latest Hugging face transformers and we advise you to use the latest version of transformers. With transformers<4.37.0, you will encounter the following error: Quickstart Here provides a code snippet with applychattemplate to show you how to load the tokenizer and model and how to generate contents. Processing Long Texts The current config.json is set for context length up to 32,768 tokens. To handle extensive inputs exceeding 32,768 tokens, we utilize YaRN, a technique for enhancing model

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

MakerQwen
TypeLanguage models
Parameters (lead)7.6B
Context32k tokens
Variants8
Runs withtransformers
Based onQwen/Qwen2.5-7B
Released2024-09-16
Popularity10.8M downloads / month
Likes1,568
LicenceCommercial licence needed

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
Qwen2.5-7B-Instruct7.6BBF16~17.5 GBWeights ↗
Qwen2.5-1.5B-Instruct1.5BBF16~3.6 GBWeights ↗
Qwen2.5-3B-Instruct3.1BBF16~7.1 GBWeights ↗
Qwen2.5-0.5B-Instruct494MBF16~1.1 GBWeights ↗
Qwen2.5-7B-Instruct-AWQ7.6BAWQWeights ↗
Qwen2.5-14B-Instruct14.8BBF16~34 GBWeights ↗
Qwen2.5-14B-Instruct-AWQ14.8BAWQWeights ↗
Qwen2.5-32B-Instruct32.8BBF16~75.4 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys qwen2-5 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (qwen2-5 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen2-5","messages":[{"role":"user","content":"Hello"}]}'

Details

Languages, data & research

Languages

en

Tags

transformers safetensors qwen2 text-generation chat conversational en eval-results text-generation-inference endpoints_compatible deploy:sagemaker deploy:azure 4-bit awq

Papers

Licence

Commercial licence needed

The weights are open but apache-2.0 needs a commercial agreement for business use. AxForge can arrange that licence and host the model for you — you pay AxForge, we settle with the model’s maker. Ask us for a quote. Read the licence ↗

Sources

Weights & code

Want Qwen2.5 on EU-owned hardware?

Request a licence + hosting quote See what’s served now

Explore

More language models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms