Model reference · open weights

sarashina2.2-tts

Available as managed deployment Audio sbintuitions Text→speech 1 variants 7k dl/mo

sarashina2.2-tts is an open-weight audio or speech model from sbintuitions. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released bysbintuitions
TypeAudio & music
TaskText→speech
Parameters (lead)810M
Context8k tokens
Based onsbintuitions/sarashina2.2-0.5b-instruct-v0.1
Released2026-04-16
Popularity7k downloads / month
LicenceUnknown

About

What sarashina2.2-tts is

Sarashina2.2-TTS is a Japanese-centric text-to-speech system built on a large language model, developed by SB Intuitions. It supports Japanese and English, delivering high pronunciation accuracy, naturalness, and stability across diverse speaking styles, with zero-shot voice generation support.

Read the full model card

Highlights

  • 🇯🇵 Japanese-Centric: Designed and optimized specifically for Japanese, with broad coverage of real-world use cases.
  • 🎯 High Accuracy: Delivers strong pronunciation accuracy on Japanese text through large-scale end-to-end training.
  • 🔒 Responsibly Sourced Training Data: Trained exclusively on legitimately acquired and properly licensed speech data.
  • 🎙️ Zero-shot Voice Generation: Reproduces a speaker's voice, speaking style, and acoustic characteristics from a short reference clip.
  • 🔊 Natural & Expressive: Produces highly natural speech with consistent quality, supporting a wide range of speaking styles including narration, broadcast, conversation, and customer service.
  • 🌐 Bilingual: Supports both Japanese and English text-to-speech synthesis.

Training Data

This model was trained on audio data collected from legitimately purchased audio sources, public speech archives, and data gathered in compliance with applicable domestic laws. During collection, we adhered to robots.txt directives and terms of service to ensure proper data acquisition.

Usage

For installation instructions, Docker setup, and detailed usage, please refer to the GitHub repository.

Audio Samples

The samples below demonstrate key capabilities of Sarashina2.2-TTS:

  • Speaking Style Variety: Transfers diverse speaking styles — narration, broadcast, conversation, customer service, and more — from reference audio.
  • Zero-shot Voice Cloning: Reproduces a speaker's voice from just a few seconds of reference speech, with no fine-tuning required.
  • Cross-lingual Generation: Preserves speaker identity and speaking style consistently across Japanese and English.
  • Code Switching: Handles mixed Japanese-English sentences naturally within a single utterance.
Quick Sample
Speaking Style Variety
Zero-shot Voice Generation
Cross-lingual Zero-shot
Code Switching

Acknowledgments

This model is built upon or incorporates code and models from the following open-source projects:

Citation

@misc{liu2026sarashina22ttstacklingkanjipolyphony,
      title={Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis},
      author={Lianbo Liu and Shiao Zhu and Kai Washizaki and Reo Yoneyama and Haesung Jeon and Mengjie Zhao and Yusuke Fujita and Hao Shi and Nao Yoshida and Yuan Gao and Roman Koshkin and Yukiya Hono and Yui Sudo},
      year={2026},
      eprint={2606.25369},
      archivePrefix={arXiv},
      primaryClass={cs.SD},
      url={https://arxiv.org/abs/2606.25369},
}

License

This model is licensed under Sarashina Model NonCommercial License Agreement.

If you are interested in using this model for commercial purposes, please feel free to contact us through our contact page.

The audio provided on this page are for research purposes only and may not be redistributed or used for commercial purposes.

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys sarashina2-2-tts for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (sarashina2-2-tts below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="sarashina2-2-tts" -F file=@audio.mp3

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms