Model reference · open weights

Fun-CosyVoice3-2512

Fun-CosyVoice3-2512 is an open-weight audio or speech model from FunAudioLLM, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio FunAudioLLM 1 variants 52k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What Fun-CosyVoice3-2512 is

👉🏻 CosyVoice 👈🏻 Fun-CosyVoice 3.0: Demos; Paper; Modelscope; Huggingface; CV3-Eval CosyVoice 2.0: Demos; Paper; Modelscope; HuggingFace CosyVoice 1.0: Demos; Paper; Modelscope; HuggingFace Highlight🔥 Fun-CosyVoice 3.0 is an advanced text-to-speech (TTS) system based on large language models (LLM), surpassing its predecessor (CosyVoice 2.0) in content consistency, speaker similarity, and prosody naturalness. It is designed for zero-shot multilingual speech synthesis in the wild. Key Features - Language Coverage: Covers 9 common languages (Chinese, English, Japanese, Korean, German, Spanish, French, Italian, Russian), 18+ Chinese dialects/accents (Guangdong, Minnan, Sichuan, Dongbei, Shan3xi, Shan1xi, Shanghai, Tianjin, Shandong, Ningxia, Gansu, etc.) and meanwhile supports both multi-lingual/cross-lingual zero-shot voice cloning. - Content Consistency & Naturalness: Achieves state-of-the-art performance in content consistency, speaker similarity, and prosody naturalness. - Pronunciation Inpainting: Supports pronunciation inpainting of Chinese Pinyin and English CMU phonemes, providing more controllability and thus suitable for production use. - Text Normalization: Supports reading of numbers, special symbols and various text formats without a traditional frontend module. - Bi-Streaming: Support both text-in streaming and audio-out streaming, and achieves latency as low as 150ms while maintaining high-quality audio output. - Instruct Support: Supports various instructions such as languages, dialects, emotions, speed, volume, etc. Roadmap - [x] 2025/12 - [x] release Fun-CosyVoice3-0.5B-2512 base model, rl model and its training/inference script - [x] release Fun-CosyVoice3-0.5B modelscope gradio space - [x] 2025/08 - [x] Thanks to the contribution from NVIDIA Yuekai Zhang, add triton trtllm runtime support and cosyvoice2 grpo training support - [x] 2025/07 - [x] release Fun-CosyVoice 3.0 eval set - [x] 2025/05 - [x] add CosyVoice2-0.5B vllm support - [x] 2024/12 - [x] 25hz CosyVoice2-0.5B released - [x] 2024/09 - [x] 25hz CosyVoice-300M base model - [x] 25hz CosyVoice-300M voice conversion function - [x] 2024/08 - [x] Repetition Aware Sampling(RAS) inference for ll

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

MakerFunAudioLLM
TypeAudio & music
Variants1
Released2025-12-11
Popularity52k downloads / month
Likes630
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
Fun-CosyVoice3-0.5B-2512BF16Weights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys fun-cosyvoice3-2512 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (fun-cosyvoice3-2512 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="fun-cosyvoice3-2512" -F file=@audio.mp3

Details

Languages, data & research

Languages

zh en fr es ja ko it ru de

Tags

onnx safetensors text-to-speech zh en fr es ja ko it ru de

Papers

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want Fun-CosyVoice3-2512 on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms