Model reference · open weights

Chatterbox-Multilingual-es-es

Available as managed deployment Audio ResembleAI Text→speech 1 variants 0 dl/mo

Chatterbox-Multilingual-es-es is an open-weight audio or speech model from ResembleAI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

MakerResembleAI
TypeAudio & music
TaskText→speech
Runs withchatterbox
Based onResembleAI/chatterbox
Released2026-04-22
Popularity0 downloads / month
LicenceOpen weights

About

What Chatterbox-Multilingual-es-es is

🎙️ Live demo: Try this model in the ResembleAI/Chatterbox-Multilingual-TTS-es-es Space.

Chatterbox Multilingual: Spanish (Spain)

Chatterbox Multilingual: Spanish (Spain) is a dedicated single-language finetune in the Chatterbox Multilingual V3 Single Language Pack. It is optimized for Spanish as spoken in Spain, with language- and region-specific behavior for expressive text-to-speech and voice cloning.

Use this model when you want tighter Spanish (Spain) quality control than the broad multilingual checkpoint. For a single model that covers all supported languages, use ResembleAI/chatterbox.

Demo

Try the hosted demo Space: ResembleAI/Chatterbox-Multilingual-TTS-es-es.

Files

  • t3_es_es.safetensors: T3 state dict in safetensors format.
  • s3gen_v3.pt / s3gen_v3.safetensors: V3 S3Gen speech decoder checkpoint.
  • grapheme_mtl_merged_expanded_v1.json: multilingual tokenizer config.

Language

  • Locale: es-ES
  • Chatterbox language ID: es

Checkpoint Metadata

  • Source step: 135500
  • Source checkpoint: t3_135500.pth.tar
  • Tensor count: 292
  • Dtype: float32
  • Text embedding shape: (2454, 1024)
  • Speech embedding shape: (8194, 1024)
  • Size: 2143990264 bytes
  • SHA256: d85844b13ea8cb45e95b8d84a55bcfeccb2d743035cf304ee3d778fc6be39546

Loader Notes

This repository contains Chatterbox Multilingual V3 single-language assets used by the linked demo Space. The T3 checkpoint is loaded with multilingual vocabulary shape 2454 and S3 speech vocabulary shape 8194.

The demo combines these model-specific assets with the shared Chatterbox inference code and companion assets needed for end-to-end speech generation.

From the published model card. Full card on the HuggingFace links in the sidebar.

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys chatterbox-multilingual-es-es for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (chatterbox-multilingual-es-es below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="chatterbox-multilingual-es-es" -F file=@audio.mp3

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms