Model reference · open weights
dots.tts-soar is an open-weight audio or speech model from dots-studio. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | dots-studio |
|---|---|
| Type | Audio & music |
| Task | Text→speech |
| Parameters (lead) | 2.2B |
| Runs with | dots_tts |
| Based on | dots-studio/dots.tts-base |
| Released | 2026-06-04 |
| Popularity | 3k downloads / month |
| Licence | Open weights |
About
dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive (AR) text-to-speech system. The backbone pairs a semantic encoder, an LLM, and an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE — no discrete codec tokens anywhere in the pipeline.
This repository hosts dots.tts-soar — the pretrained backbone further refined with Self-corrective Alignment (SCA), a reward-free flow-matching-native post-training stage. SCA pushes the model to the highest zero-shot fidelity and speaker similarity of the three releases and is the recommended default for production zero-shot voice cloning.
conda create -n dots_tts python=3.10 -y
conda activate dots_tts
python -m pip install --upgrade pip
python -m pip install "git+https://github.com/studio-dots-ai/dots.tts.git" \
-c "https://raw.githubusercontent.com/studio-dots-ai/dots.tts/main/constraints/recommended.txt"
# Continuation voice cloning (reference audio + transcript) — recommended
dots.tts \
--model-name-or-path dots-studio/dots.tts-soar \
--text "Hello, this is a zero-shot voice cloning demonstration." \
--prompt-audio /path/to/reference.wav \
--prompt-text "The exact transcript of the reference audio." \
--output clone.wav
from dots_tts.runtime import DotsTtsRuntime
import soundfile as sf
runtime = DotsTtsRuntime.from_pretrained(
"dots-studio/dots.tts-soar",
precision="bfloat16",
)
result = runtime.generate(
text="Hello, this is a quick speech synthesis test.",
prompt_audio_path="/path/to/reference.wav",
prompt_text="The exact transcript of the reference audio.",
num_steps=10,
guidance_scale=1.2,
)
sf.write("output.wav", result["audio"].float().cpu().squeeze().numpy(), result["sample_rate"])
| Flag | Recommended | Notes |
|---|---|---|
--num-steps | 10–32 | Flow-matching sampling steps; higher = better quality, slower |
--guidance-scale | 1.2 (default) | Standard CFG; SCA already tightens text and timbre adherence so small CFG suffices |
Both dots.tts-base and dots.tts-soar are valid fine-tuning starting points. Pick dots.tts-soar when you want to inherit its tightened text/timbre alignment on top of the pretrained backbone. See the training script and smoke config in the source repository:
accelerate launch scripts/train_dots_tts.py --config configs/dots_tts.yaml
A frozen AudioVAE encodes 48 kHz mono waveform into a continuous latent and decodes it back via a BigVGAN-style causal decoder. An autoregressive backbone predicts that latent one patch at a time:
Self-corrective Alignment is a reward-free, flow-matching-native post-training stage applied on top of dots.tts-base. It improves text and speaker adherence without changing inference cost or sampling schedule.
dots.tts-soar| Model | Params | test-en WER↓ / SIM↑ | test-zh WER↓ / SIM↑ | test-zh-hard WER↓ / SIM↑ | Avg WER↓ / SIM↑ |
|---|---|---|---|---|---|
| Seed-TTS | — | 2.25 / 76.2 | 1.12 / 79.6 | 7.59 / 77.6 | 3.65 / 77.8 |
| Qwen3-TTS | 1.7B | 1.23 / 71.7 | 1.22 / 77.0 | 6.76 / 74.8 | 3.07 / 74.5 |
| VoxCPM 2 | 2B | 1.84 / 75.3 | 0.97 / 79.5 | 8.13 / 75.3 | 3.65 / 76.7 |
| dots.tts-base | 2B | 1.34 / 76.8 | 0.96 / 80.5 | 6.46 / 79.2 | 2.92 / 78.8 |
| dots.tts-soar | 2B | 1.30 / 77.1 | 0.94 / 81.0 | 6.60 / 79.5 | 2.95 / 79.2 |
| Model | Avg WER↓ | Avg SIM↑ |
|---|---|---|
| MiniMax | 2.8 | 76.6 |
| Fish-Audio S2 | 3.7 | 78.0 |
| VoxCPM 2 | 5.7 | 82.3 |
| dots.tts-base | 6.6 | 83.5 |
| dots.tts-soar | 6.8 | 83.9 |
| Model | en→zh SIM↑ | zh→en SIM↑ |
|---|---|---|
| CosyVoice 3 (1.5B) | 66.9 | 66.4 |
| dots.tts-base | 74.6 | 71.9 |
| dots.tts-soar | 75.0 | 72.8 |
On head-to-head judging vs. gpt-4o-mini-tts, dots.tts-soar posts 65.7% on Syntactic Complexity — above every closed-source system listed, while keeping competitive Emotions / Questions scores.
See the project README for full benchmark tables.
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys dots-tts-soar for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (dots-tts-soar below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \ -H "Authorization: Bearer $AXFORGE_API_KEY" \ -F model="dots-tts-soar" -F file=@audio.mp3
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.