Model reference · open weights

Qwen3-TTS-12Hz-VoiceDesign

Qwen3-TTS-12Hz-VoiceDesign is an open-weight audio or speech model from Qwen, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio Qwen 1 variants 285k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What Qwen3-TTS-12Hz-VoiceDesign is

Qwen3-TTS &nbsp&nbsp🤗 <a href="https://huggingface.co/collections/Qwen/qwen3-tts"Hugging Face</a&nbsp&nbsp | &nbsp&nbsp🤖 <a href="https://modelscope.cn/collections/Qwen/Qwen3-TTS"ModelScope</a&nbsp&nbsp | &nbsp&nbsp📑 <a href="https://qwen.ai/blog?id=qwen3tts-0115"Blog</a&nbsp&nbsp | &nbsp&nbsp📑 <a href="https://huggingface.co/papers/2601.15621"Paper</a&nbsp&nbsp | &nbsp&nbsp💻 <a href="https://github.com/QwenLM/Qwen3-TTS"GitHub</a We release Qwen3-TTS, a series of powerful speech generation models developed by Qwen, offering comprehensive support for voice cloning, voice design, ultra-high-quality human-like speech generation, and natural language-based voice control. Overview Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles. Key features: Powerful Speech Representation: Powered by the self-developed Qwen3-TTS-Tokenizer-12Hz, it achieves efficient acoustic compression and high-dimensional semantic modeling. Universal End-to-End Architecture: Utilizing a discrete multi-codebook LM architecture to bypass traditional information bottlenecks. Extreme Low-Latency Streaming Generation: Supports streaming generation with end-to-end synthesis latency as low as 97ms. Intelligent Voice Control: Supports speech generation driven by natural language instructions for flexible control over timbre, emotion, and prosody. Quickstart Environment Setup Install the qwen-tts Python package from PyPI: Python Package Usage Evaluation Zero-shot speech generation on the Seed-TTS test set (Word Error Rate (WER, ↓)): Citation If you find our paper and code useful in your research, please consider giving a star ⭐ and citation 📝:

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

MakerQwen
TypeAudio & music
Parameters (lead)1.9B
Variants1
Runs withqwen-tts
Released2026-01-21
Popularity285k downloads / month
Likes398
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
Qwen3-TTS-12Hz-1.7B-VoiceDesign1.9BBF16~4.4 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys qwen3-tts-12hz-voicedesign for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (qwen3-tts-12hz-voicedesign below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="qwen3-tts-12hz-voicedesign" -F file=@audio.mp3

Details

Languages, data & research

Tags

qwen-tts safetensors qwen3_tts audio tts qwen multilingual text-to-speech

Papers

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want Qwen3-TTS-12Hz-VoiceDesign on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms