Model reference · open weights

kitten-tts-2

NEW · this week Audio KittenML Text→speech 1 build Its own licence terms 546 dl/mo

kitten-tts-2 is an open-weight audio or speech model from KittenML. kitten-tts-2 (BF16) weighs 1.0 GB; the smallest configuration that runs it is RTX 3060 12 GB.

  • Kitten-tts-2 is a 1.7B parameter text-to-speech model by KittenML that generates 24 kHz audio with in-context voice cloning.
  • It supports 47 voices, including nine specific language options for Arabic, Chinese, French, German, Hindi, Italian, Portuguese, Russian, and Spanish.
  • The model is released under the Stellon Labs Community License, which allows free research and limited commercial use.

Summary of the KittenML/kitten-tts-2 model card, 2026-10-05

What it is

Released byKittenML
Released2026-09-30
VRAM1.0 GB for the weights

What it runs on

Memory and cards for kitten-tts-2 (BF16)

1.0 GBweights, file size
1.6 GBruntime overhead
CardOne streamMemory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

From the model card

What KittenML says about kitten-tts-2

Read the model card

A 1.7B speech language model with in-context voice cloning. It reads text and writes S3 codec tokens, which a vocoder turns into 24 kHz audio.

pip install kittenml
from kittenml import KittenTTS
import soundfile as sf

m = KittenTTS("KittenML/kitten-tts-2")

# a built-in voice
audio = m.generate("One day, a little girl named Lily found a needle in her room.",
                   voice="Bruno")

# or clone one, from 5-30 seconds of a single speaker
audio = m.generate("This is my own voice.", reference="my_voice.wav")

sf.write("output.wav", audio, m.sample_rate)

Everything the model needs is in this repository, so no Hugging Face login is required.

Voices

m.available_voices lists all 47. Bella, Jasper, Luna, Bruno, Rosie, Hugo, Kiki and Leo are the same speakers as in KittenTTS 0.8, so existing code keeps working.

Nine are named after a language rather than a person — Arabic, Chinese, French, German, Hindi, Italian, Portuguese, Russian, Spanish — and are how you reach those languages, since the voice is what carries the accent:

m.generate("Guten Morgen. Ich wünsche dir einen wunderschönen Tag.",
           voice="German", normalize=False)

Pass normalize=False for non-English text: the text normalizer is English-tuned and will mangle numbers and dates in other languages.

Expression controls

Beta. Emotion control steers delivery rather than guaranteeing it, and the effect varies by voice and by sentence.

m.generate("[joyful] We won the grant  I can (((hardly))) believe it!",
           voice="Kiki", preset="expressive")

A leading [emotion] tag, inline `` tags and (((emphasis))) spans reach the model as markup rather than being spoken, and switch on its expression conditioning.

Emotions — one leading tag sets the emotion for the whole line:

[angry] [contemplative] [excited] [joyful] [mundane] [nervous] [sad] [stern] [surprised] [tender]

Vocal events — inline, anywhere in the line:

Emphasis — triple parentheses stress a word or short phrase:

m.generate("I told you (((never))) to open that door.", voice="Victor")

Only these twenty tags are recognised. They are the most common of the many in the training data, so a rarer one such as [reverent] is spoken as ordinary text rather than treated as markup — as is anything else bracketed, like section [3] or x < 5.

Weight variants

The language model ships in three packings of the same weights. weights= picks one; model.available_weights lists them.

weights=On disk
"packed" (default)954 MB1.58-bit ternary body, bf16 embedding. Lossless
"emb4"469 MBSame body in TL2, plus a 4-bit embedding. Lossy
"full"3469 MBPlain bf16
m = KittenTTS("KittenML/kitten-tts-2", weights="emb4")

Only the variant you ask for is downloaded.

emb4 halves the download, and the saving is almost entirely the token embedding — 324M parameters that packed has to leave at bf16 because they are not ternary. Its transformer body is still bit-exact; the embedding is not, at L2 relative error 0.118 against bf16. Measured at export, that costs roughly a tenth of a point of perplexity on internal evaluations. Small, but it is the one lossy thing here, which is why packed stays the default.

Decoders

Audio is decoded in two stages, and the first can be swapped for a smaller distilled student with weights packed to 4 or 8 bits:

m = KittenTTS("KittenML/kitten-tts-2", decoder="student_w4")
DecoderFlow on disk
default459 MBBest quality
student_w439 MBDistilled single-step student, weights packed to 4 bits

These trade fidelity for footprint, not for speed: quantisation shrinks storage and memory bandwidth, not arithmetic.

What is in here

lm/          the speech language model, plus its spk_proj speaker head
speaker/     the speaker-embedding model used when cloning
voices/      reference clips, transcripts, and precomputed embeddings
decoders/    the optional 4-bit decoder
cpp/         GGUF weights for the llama.cpp fork, see "Running on CPU"
config.json  token layout, decode presets, voice and decoder indexes

The weights are 947 MiB: 910 MiB for the language model and 37 MiB for the decoder. A load pulls those rather than the whole repository, and only the decoder you ask for.

The language model's linear weights are ternary — within every 128-wide group each value is exactly one of {-scale, 0, +scale} — so bf16 spends 16 bits to say one of three things. lm/model-ternary.safetensors packs them five trits to a byte, 1.6 bits per weight, with each group keeping its scale at full precision:

lm/model.safetensors3.47 GB, bf16 throughout
lm/model-ternary.safetensors0.95 GB, the same weights, 3.6x smaller

The packing is exact rather than approximate, so the two produce bit-identical audio. config.json points lm_packed at the smaller file and that is what gets downloaded; the full file stays for anything loading this with plain transformers.

Running on CPU

KittenTTS 2 runs on CPU out of the box — device is auto-detected — but the fastest way is kitten-tts-2-cpp, our llama.cpp fork. The cpp/ directory here holds what it needs: the GGUF language model and its decoders.

Requirements

Python 3.10 or later, and PyTorch.

License

These weights are released under the Stellon Labs Community License. Research, non-commercial and limited commercial use are free of charge; the commercial grant ends once you or your affiliates pass USD $1,000,000 in annual revenue or in total cumulative funding, at which point you need a separate license from Stellon Labs. Distributing the weights, a derivative, or a product built on them carries attr

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms