Model reference · open weights

supertonic-3

Available as managed deployment Audio supertone-oss-archive Text→speech 1 variants 647 dl/mo

supertonic-3 is an open-weight audio or speech model from supertone-oss-archive. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released bysupertone-oss-archive
TypeAudio & music
TaskText→speech
Runs withsupertonic
Released2026-09-08
Popularity647 downloads / month
LicenceOpen weights

About

What supertonic-3 is

Supertonic is a lightweight text-to-speech system for local inference. It runs with ONNX Runtime entirely on your device, with no cloud call required for synthesis.

Supertonic 3 expands the open-weight release from 5 to 31 languages, improves reading stability, and reduces repeat/skip failures.

Read the full model card

Quick Start

Install the Python SDK and generate speech immediately. On first run, the SDK downloads the model assets from Hugging Face.

pip install supertonic
from supertonic import TTS

tts = TTS(auto_download=True)
style = tts.get_voice_style(voice_name="M1")

text = "A gentle breeze moved through the open window while everyone listened to the story."
wav, duration = tts.synthesize(text, voice_style=style, lang="en")

tts.save_audio(wav, "output.wav")
print(f"Generated {duration:.2f}s of audio")

What's New in Supertonic 3

  • 31 languages: expanded from the 5-language Supertonic 2 release.
  • More stable reading: fewer repeat and skip failures, especially on short and long utterances.
  • Higher speaker similarity: improved similarity across the shared-language set compared with Supertonic 2.
  • Expression tags: supports simple tags such as , , and ``.

Custom Voices and Audio Samples

The open-weight package includes fixed preset voice styles for immediate local inference. If you want to hear how Supertonic 3 performs with zero-shot custom voice styles, visit the Audio Sample Demo to compare reference audio and generated speech across several use cases. To create your own Supertonic 3 voice-style JSON from reference audio, use Supertonic Voice Builder; purchased Voice Builder styles include downloadable embeddings for both Supertonic 2 and Supertonic 3.

Here are a few reference/generated pairs from the audio sample demo:

Call center, English Text: Good morning, thank you for calling. How can I help you today?

Reference voiceSupertonic 3 output

Character voice, Japanese Text: ふふっ、退屈してたところなの。ちょうどいい遊び相手、見つけたかも♪

Reference voiceSupertonic 3 output

Elder character voice, Korean Text: 혼자 떠나기엔 길이 험하구나. 이 낡은 검을 가져가거라. 언젠가 어둠이 네 이름을 부르더라도, 부디 빛을 잊지 말거라.

Reference voiceSupertonic 3 output

Audiobook, English Text: I was not afraid of silence. I had lived with it long enough to know that, sometimes, it speaks more honestly than people do.

Reference voiceSupertonic 3 output

Audiobook, Japanese Text: その朝、ロンドンの霧はいつになく低く垂れこめていた。私はただの訪問者だと思っていたが、ホームズの目はすでに別の結論にたどり着いていた。

Reference voiceSupertonic 3 output

News, English Text: Here’s a story worth paying attention to. Supertone has released Supertonic 3, its on-device TTS model. This version expands support to thirty-one languages and improves reading stability.

Reference voiceSupertonic 3 output

Performance Highlights

Supertonic 3 is designed for practical on-device inference: compact enough to run locally, while staying competitive with much larger open TTS systems.

Reading Accuracy

Across measured languages, Supertonic 3 stays within a competitive WER/CER range against much larger open TTS models such as VoxCPM2, while preserving a lightweight on-device deployment path. Asterisked languages use CER; the others use WER.

Supertonic 2 to Supertonic 3

Compared with Supertonic 2, Supertonic 3 reduces repeat and skip failures, improves speaker similarity across the shared-language set, and expands language coverage from 5 to 31 languages.

Runtime Footprint

Supertonic 3 runs fast on CPU, even compared with larger baselines measured on A100 GPU, and uses substantially less memory. It does not require a GPU, which makes local, browser, and edge deployment much easier.

Model Size

At about 99M parameters across the public ONNX assets, Supertonic 3 is much smaller than 0.7B to 2B class open TTS systems. The smaller model size is a practical advantage for download size, startup time, and on-device inference.

Supported Languages

CodeLanguageCodeLanguageCodeLanguageCodeLanguage
enEnglishkoKoreanjaJapanesearArabic
bgBulgariancsCzechdaDanishdeGerman
elGreekesSpanishetEstonianfiFinnish
frFrenchhiHindihrCroatianhuHungarian
idIndonesianitItalianltLithuanianlvLatvian
nlDutchplPolishptPortugueseroRomanian
ruRussianskSlovakslSloveniansvSwedish
trTurkishukUkrainianviVietnamese

License

This project's sample code is released under the MIT License. See the GitHub repository for details.

The accompanying model is released under the OpenRAIL-M License. See the LICENSE file in this repository for details.

This model was trained using PyTorch, which is licensed under the BSD 3-Clause License but is not redistributed with this project. See the PyTorch license for details.

Copyright (c) 2026 Supertone Inc.

From the published model card. Full card on the HuggingFace links in the sidebar.

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys supertone-oss-archive-supertonic-3 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (supertone-oss-archive-supertonic-3 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="supertone-oss-archive-supertonic-3" -F file=@audio.mp3

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms