Model reference · open weights
Qwen2.5-Omni is an open-weight language model from unsloth. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Maker | unsloth |
|---|---|
| Type | Language models |
| Task | Omni (any→any) |
| Runs with | transformers |
| Based on | Qwen/Qwen2.5-Omni-7B |
| Released | 2025-05-28 |
| Popularity | 18k downloads / month |
| Licence | Commercial licence needed |
About
Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner.
Omni and Novel Architecture: We propose Thinker-Talker architecture, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. We propose a novel position embedding, named TMRoPE (Time-aligned Multimodal RoPE), to synchronize the timestamps of video inputs with audio.
Real-Time Voice and Video Chat: Architecture designed for fully real-time interactions, supporting chunked input and immediate output.
Natural and Robust Speech Generation: Surpassing many existing streaming and non-streaming alternatives, demonstrating superior robustness and naturalness in speech generation.
Strong Performance Across Modalities: Exhibiting exceptional performance across all modalities when benchmarked against similarly sized single-modality models. Qwen2.5-Omni outperforms the similarly sized Qwen2-Audio in audio capabilities and achieves comparable performance to Qwen2.5-VL-7B.
Excellent End-to-End Speech Instruction Following: Qwen2.5-Omni shows performance in end-to-end speech instruction following that rivals its effectiveness with text inputs, evidenced by benchmarks such as MMLU and GSM8K.
We conducted a comprehensive evaluation of Qwen2.5-Omni, which demonstrates strong performance across all modalities when compared to similarly sized single-modality models and closed-source models like Qwen2.5-VL-7B, Qwen2-Audio, and Gemini-1.5-pro. In tasks requiring the integration of multiple modalities, such as OmniBench, Qwen2.5-Omni achieves state-of-the-art performance. Furthermore, in single-modality tasks, it excels in areas including speech recognition (Common Voice), translation (CoVoST2), audio understanding (MMAU), image reasoning (MMMU, MMStar), video understanding (MVBench), and speech generation (Seed-tts-eval and subjective naturalness).
| Dataset | Qwen2.5-Omni-7B | Qwen2.5-Omni-3B | Other Best | Qwen2.5-VL-7B | GPT-4o-mini |
|---|---|---|---|---|---|
| MMMUval | 59.2 | 53.1 | 53.9 | 58.6 | 60.0 |
| MMMU-Prooverall | 36.6 | 29.7 | - | 38.3 | 37.6 |
| MathVistatestmini | 67.9 | 59.4 | 71.9 | 68.2 | 52.5 |
| MathVisionfull | 25.0 | 20.8 | 23.1 | 25.1 | - |
| MMBench-V1.1-ENtest | 81.8 | 77.8 | 80.5 | 82.6 | 76.0 |
| MMVetturbo | 66.8 | 62.1 | 67.5 | 67.1 | 66.9 |
| MMStar | 64.0 | 55.7 | 64.0 | 63.9 | 54.8 |
| MMEsum | 2340 | 2117 | 2372 | 2347 | 2003 |
| MuirBench | 59.2 | 48.0 | - | 59.2 | - |
| CRPErelation | 76.5 | 73.7 | - | 76.4 | - |
| RealWorldQAavg | 70.3 | 62.6 | 71.9 | 68.5 | - |
| MME-RealWorlden | 61.6 | 55.6 | - | 57.4 | - |
| MM-MT-Bench | 6.0 | 5.0 | - | 6.3 | - |
| AI2D | 83.2 | 79.5 | 85.8 | 83.9 | - |
| TextVQAval | 84.4 | 79.8 | 83.2 | 84.9 | - |
| DocVQAtest | 95.2 | 93.3 | 93.5 | 95.7 | - |
| ChartQAtest Avg | 85.3 | 82.8 | 84.9 | 87.3 | - |
| OCRBench_V2en | 57.8 | 51.7 | - | 56.3 | - |
| Dataset | Qwen2.5-Omni-7B | Qwen2.5-Omni-3B | Qwen2.5-VL-7B | Grounding DINO | Gemini 1.5 Pro |
|---|---|---|---|---|---|
| Refcocoval | 90.5 | 88.7 | 90.0 | 90.6 | 73.2 |
| RefcocotextA | 93.5 | 91.8 | 92.5 | 93.2 | 72.9 |
| RefcocotextB | 86.6 | 84.0 | 85.4 | 88.2 | 74.6 |
| Refcoco+val | 85.4 | 81.1 | 84.2 | 88.2 | 62.5 |
| Refcoco+textA | 91.0 | 87.5 | 89.1 | 89.0 | 63.9 |
| Refcoco+textB | 79.3 | 73.2 | 76.9 | 75.9 | 65.0 |
| Refcocog+val | 87.4 | 85.0 | 87.2 | 86.1 | 75.2 |
| Refcocog+test | 87.9 | 85.1 | 87.2 | 87.0 | 76.2 |
| ODinW | 42.4 | 39.2 | 37.3 | 55.0 | 36.7 |
| PointGrounding | 66.5 | 46.2 | 67.3 | - | - |
| Dataset | Qwen2.5-Omni-7B | Qwen2.5-Omni-3B | Other Best | Qwen2.5-VL-7B | GPT-4o-mini |
|---|---|---|---|---|---|
| Video-MMEw/o sub | 64.3 | 62.0 | 63.9 | 65.1 | 64.8 |
| Video-MMEw sub | 72.4 | 68.6 | 67.9 | 71.6 | - |
| MVBench | 70.3 | 68.7 | 67.2 | 69.6 | - |
| EgoSchematest | 68.6 | 61 |
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys unsloth-qwen2-5-omni for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (unsloth-qwen2-5-omni below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"unsloth-qwen2-5-omni","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.