Model reference · open weights

MeloTTS-Chinese

MeloTTS-Chinese is an open-weight audio or speech model from myshell-ai, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio myshell-ai 1 variants 68k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What MeloTTS-Chinese is

MeloTTS MeloTTS is a high-quality multi-lingual text-to-speech library by MyShell.ai. Supported languages include: Some other features include: - The Chinese speaker supports mixed Chinese and English. - Fast enough for CPU real-time inference. Usage Without Installation An unofficial live demo is hosted on Hugging Face Spaces. Use it on MyShell There are hundreds of TTS models on MyShell, much more than MeloTTS. See examples here. More can be found at the widget center of MyShell.ai. Install and Use Locally Follow the installation steps here before using the following snippet: Join the Community Open Source AI Grant We are actively sponsoring open-source AI projects. The sponsorship includes GPU resources, fundings and intellectual support (collaboration with top research labs). We welcome both reseach and engineering projects, as long as the open-source community needs them. Please contact Zengyi Qin if you are interested. Contributing If you find this work useful, please consider contributing to the GitHub repo. - Many thanks to @fakerybakery for adding the Web UI and CLI part. License This library is under MIT License, which means it is free for both commercial and non-commercial use. Acknowledgements This implementation is based on TTS, VITS, VITS2 and Bert-VITS2. We appreciate their awesome work.

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makermyshell-ai
TypeAudio & music
Variants1
Runs withtransformers
Released2024-02-29
Popularity68k downloads / month
Likes103
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
MeloTTS-ChineseBF16Weights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys melotts-chinese for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (melotts-chinese below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="melotts-chinese" -F file=@audio.mp3

Details

Languages, data & research

Languages

ko

Tags

transformers text-to-speech ko endpoints_compatible

Licence

Open weights

Open weights under mit — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want MeloTTS-Chinese on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms