Model reference · open weights

Zonos-hybrid

Available as managed deployment Audio Zyphra Text→speech 1 variants 1k dl/mo

Zonos-hybrid is an open-weight audio or speech model from Zyphra. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byZyphra
TypeAudio & music
TaskText→speech
Parameters (lead)1.7B
Runs withzonos
Released2025-02-06
Popularity1k downloads / month
LicenceOpen weights

About

What Zonos-hybrid is

alt="Title card" style="width: 500px; height: auto; object-position: center top;">


Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers.

Our model enables highly natural speech generation from text prompts when given a speaker embedding or audio prefix, and can accurately perform speech cloning when given a reference clip spanning just a few seconds. The conditioning setup also allows for fine control over speaking rate, pitch variation, audio quality, and emotions such as happiness, fear, sadness, and anger. The model outputs speech natively at 44kHz.

For more details and speech samples, check out our blog here
We also have a hosted version available at playground.zyphra.com/audio

Zonos follows a straightforward architecture: text normalization and phonemization via eSpeak, followed by DAC token prediction through a transformer or hybrid backbone. An overview of the architecture can be seen below.

Read the full model card
 alt="Architecture diagram"
 style="width: 1000px;
        height: auto;
        object-position: center top;">

Usage

Python

import torch
import torchaudio
from zonos.model import Zonos
from zonos.conditioning import make_cond_dict

model = Zonos.from_pretrained("Zyphra/Zonos-v0.1-hybrid", device="cuda")
# model = Zonos.from_pretrained("Zyphra/Zonos-v0.1-transformer", device="cuda")

wav, sampling_rate = torchaudio.load("assets/exampleaudio.mp3")
speaker = model.make_speaker_embedding(wav, sampling_rate)

cond_dict = make_cond_dict(text="Hello, world!", speaker=speaker, language="en-us")
conditioning = model.prepare_conditioning(cond_dict)

codes = model.generate(conditioning)

wavs = model.autoencoder.decode(codes).cpu()
torchaudio.save("sample.wav", wavs[0], model.autoencoder.sampling_rate)

Gradio interface (recommended)

uv run gradio_interface.py
# python gradio_interface.py

This should produce a sample.wav file in your project root directory.

For repeated sampling we highly recommend using the gradio interface instead, as the minimal example needs to load the model every time it is run.

Features

  • Zero-shot TTS with voice cloning: Input desired text and a 10-30s speaker sample to generate high quality TTS output
  • Audio prefix inputs: Add text plus an audio prefix for even richer speaker matching. Audio prefixes can be used to elicit behaviours such as whispering which can otherwise be challenging to replicate when cloning from speaker embeddings
  • Multilingual support: Zonos-v0.1 supports English, Japanese, Chinese, French, and German
  • Audio quality and emotion control: Zonos offers fine-grained control of many aspects of the generated audio. These include speaking rate, pitch, maximum frequency, audio quality, and various emotions such as happiness, anger, sadness, and fear.
  • Fast: our model runs with a real-time factor of ~2x on an RTX 4090
  • Gradio WebUI: Zonos comes packaged with an easy to use gradio interface to generate speech
  • Simple installation and deployment: Zonos can be installed and deployed simply using the docker file packaged with our repository.

Installation

At the moment this repository only supports Linux systems (preferably Ubuntu 22.04/24.04) with recent NVIDIA GPUs (3000-series or newer, 6GB+ VRAM).

See also Docker Installation

System dependencies

Zonos depends on the eSpeak library phonemization. You can install it on Ubuntu with the following command:

apt install -y espeak-ng
Python dependencies

We highly recommend using a recent version of uv for installation. If you don't have uv installed, you can install it via pip: pip install -U uv.

Installing into a new uv virtual environment (recommended)
uv sync
uv sync --extra compile
Installing into the system/actived environment using uv
uv pip install -e .
uv pip install -e .[compile]
Installing into the system/actived environment using pip
pip install -e .
pip install --no-build-isolation -e .[compile]
Confirm that it's working

For convenience we provide a minimal example to check that the installation works:

uv run sample.py
# python sample.py

Docker installation

git clone https://github.com/Zyphra/Zonos.git
cd Zonos

# For gradio
docker compose up

# Or for development you can do
docker build -t Zonos .
docker run -it --gpus=all --net=host -v /path/to/Zonos:/Zonos -t Zonos
cd /Zonos
python sample.py # this will generate a sample.wav in /Zonos

Citation

If you find this model useful in an academic context please cite as:

@misc{zyphra2025zonos,
  title     = {Zonos-v0.1: An Expressive, Open-Source TTS Model},
  author    = {Dario Sucic, Mohamed Osman, Gabriel Clark, Chris Warner, Beren Millidge},
  year      = {2025},
}

From the published model card. Full card on the HuggingFace links in the sidebar.

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys zonos-hybrid for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (zonos-hybrid below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="zonos-hybrid" -F file=@audio.mp3

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms