Model reference · open weights

Chatterbox-Multilingual-es-mx-latam

Available as managed deployment Audio ResembleAI Text→speech 1 variants 0 dl/mo

Chatterbox-Multilingual-es-mx-latam is an open-weight audio or speech model from ResembleAI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

MakerResembleAI
TypeAudio & music
TaskText→speech
Runs withchatterbox
Based onResembleAI/chatterbox
Released2026-04-22
Popularity0 downloads / month
LicenceOpen weights

About

What Chatterbox-Multilingual-es-mx-latam is

🎙️ Live demo: Try this model in the ResembleAI/Chatterbox-Multilingual-TTS-es-mx-latam Space.

Chatterbox Multilingual: Latin American Spanish

Chatterbox Multilingual: Latin American Spanish is a dedicated single-language finetune in the Chatterbox Multilingual V3 Single Language Pack. It is optimized for Spanish as spoken in Latin America and Mexico, with language- and region-specific behavior for expressive text-to-speech and voice cloning.

Use this model when you want tighter Latin American Spanish quality control than the broad multilingual checkpoint. For a single model that covers all supported languages, use ResembleAI/chatterbox.

Demo

Try the hosted demo Space: ResembleAI/Chatterbox-Multilingual-TTS-es-mx-latam.

Files

  • t3_es_mx_latam.safetensors: T3 state dict in safetensors format.
  • s3gen_v3.pt / s3gen_v3.safetensors: V3 S3Gen speech decoder checkpoint.
  • grapheme_mtl_merged_expanded_v1.json: multilingual tokenizer config.

Language

  • Locale: es-419 / es-MX
  • Chatterbox language ID: es

Checkpoint Metadata

  • Source step: 138500
  • Source checkpoint: t3_138500.pth.tar
  • Tensor count: 292
  • Dtype: float32
  • Text embedding shape: (2454, 1024)
  • Speech embedding shape: (8194, 1024)
  • Size: 2143990280 bytes
  • SHA256: c66c4517f11c2b35a56e28615c0689deb864cbc411e329552223bbd0a6a063f8

Loader Notes

This repository contains Chatterbox Multilingual V3 single-language assets used by the linked demo Space. The T3 checkpoint is loaded with multilingual vocabulary shape 2454 and S3 speech vocabulary shape 8194.

The demo combines these model-specific assets with the shared Chatterbox inference code and companion assets needed for end-to-end speech generation.

From the published model card. Full card on the HuggingFace links in the sidebar.

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys chatterbox-multilingual-es-mx-latam for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (chatterbox-multilingual-es-mx-latam below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="chatterbox-multilingual-es-mx-latam" -F file=@audio.mp3

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms