Model reference · open weights

rumik-oss-1

Available as managed deployment Licence fee Audio rumik-ai Text→speech 1 variants 952 dl/mo

rumik-oss-1 is an open-weight audio or speech model from rumik-ai. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byrumik-ai
TypeAudio & music
TaskText→speech
Parameters (lead)3.4B
Context8k tokens
Runs withtransformers
Based onrumik-ai/rumik-oss-1-base
Released2026-09-06
Popularity952 downloads / month
LicenceCommercial licence needed

About

What rumik-oss-1 is

rumik-oss 1 is a 3b multilingual text-to-speech model from rumik ai, trained on fewer than 70,000 hours of speech while performing competitively with existing TTS models. it combines code-switched synthesis, description-conditioned delivery, and inline vocalization control, with 24 khz audio output.

we also release rumik-oss 1 base, our final speaker-conditioned checkpoint before post-training. it supports multilingual synthesis with ira, aisha, siya, and zoya and provides a starting point for the community to research and post-train the base model from scratch.

we describe the post-training of rumik-oss 1 and the decisions behind its training curriculum in our blog. our technical report will provide an in-depth account of the architecture, data preparation, training, and evaluation.

Read the full model card

model overview

rumik-oss 1 extends tiny aya fire with discrete speech tokens from the mimi codec. following the flattened codec-token formulation used in llama-mimi, text conditioning and audio generation share a single autoregressive sequence. the model predicts eight codebook tokens for each audio frame, in codebook order, before advancing to the next frame. generated tokens are regrouped into codec frames and reconstructed as a waveform by the frozen mimi decoder.

we release the model with 4 voices:

  • Ira
  • Aisha
  • Siya
  • Zoya

unlike other TTS models, where each voice is trained specifically for one language or performs best in one language, our voices perform equally well across all 22 languages.

capabilities

  • multilingual synthesis: 22 indic languages in their native scripts and romanized forms, plus english, with support for single-language and code-switched synthesis.
  • delivery conditioning: control tone, accent, and pace through the description format for various scenarios.
  • vocalization control: inline tags for laughter, chuckles, and sighs.

delivery controls

the prefix lets users control the global tone and pace of the audio, while inline tags let you insert, , and at the intended positions.

controlvalues
speakerIra, Aisha, Siya, Zoya
tonehappy, sad, angry, excited, professional
accentHindi, Telugu, Tamil, Kannada, Bengali, Punjabi, Indian English
paceslow, fast, steady
inline vocalization, , ``

audio samples

ira · hindi · sad · slow pace

माँ, आज फिर तुम्हारे लिए चाय बना दी। <sigh> कप ठंडा हो गया, पर तुम्हारा इंतज़ार नहीं।

aisha · english · sad · slow pace

I still play your last message, Dad. You only said, call me when you get home. I got home. I just never got to tell you.

siya · telugu · excited · fast pace

అమ్మా, నాకు ఉద్యోగం వచ్చింది! నిజంగా వచ్చింది! నువ్వు నా కోసం చేసిన ప్రతి త్యాగం నాకు గుర్తుంది. <laugh> ఈ రోజు నీకు నచ్చిన స్వీట్లు నేనే కొనిస్తాను!

zoya · tamil · angry · fast pace

நீ வருவேன்னு சொன்னதால்தான் இவ்வளவு நேரம் காத்திருந்தேன்! ஒரு போன் கூட பண்ண முடியலையா? ஒவ்வொரு முறையும் மன்னிப்பு கேட்டா மட்டும் எல்லாம் சரியாகிடாது!

ira · english · professional · steady pace

Good evening, passengers. Boarding for the flight to Bengaluru will begin shortly at gate twelve. Please keep your boarding pass ready. Thank you for your patience, and have a pleasant journey.

aisha · bengali · happy · steady pace

এত দিন পরে তোমাকে দেখে কী যে ভালো লাগছে! তোমার পছন্দের সব রান্না করেছি। আজ আর কোথাও যেতে দেব না, সবাই মিলে অনেক গল্প করব।

siya · kannada · professional · steady pace

ನಮ್ಮ ಹೊಸ ಗ್ರಂಥಾಲಯಕ್ಕೆ ಸ್ವಾಗತ. ಇಲ್ಲಿ ಕನ್ನಡ ಮತ್ತು ಇಂಗ್ಲಿಷ್ ಪುಸ್ತಕಗಳು ಲಭ್ಯವಿವೆ. ಸದಸ್ಯರಾಗಲು ನಿಮ್ಮ ಹೆಸರು ಮತ್ತು ವಿಳಾಸ ನೀಡಿ. ಓದಲು ಶಾಂತವಾದ ಸ್ಥಳವೂ ಇದೆ.

zoya · punjabi · excited · fast pace

ਮਾਂ, ਮੇਰਾ ਦਾਖ਼ਲਾ ਹੋ ਗਿਆ! ਸੱਚੀਂ, ਚਿੱਠੀ ਆ ਗਈ ਹੈ! ਜਿਹੜਾ ਸੁਪਨਾ ਅਸੀਂ ਇਕੱਠੇ ਵੇਖਿਆ ਸੀ, ਉਹ ਅੱਜ ਪੂਰਾ ਹੋ ਗਿਆ। ਹੁਣ ਸਭ ਨੂੰ ਫ਼ੋਨ ਕਰ!

ira · hindi + english · happy · steady pace

आज की meeting में सबको हमारा idea पसंद आया। I was so nervous, लेकिन तुमने कहा था ना, बस दिल से बोलो। <chuckle> अब coffee मेरी तरफ़ से, और cake तुम्हारी तरफ़ से!

aisha · telugu + english · professional · steady pace

మన కొత్త demo సిద్ధంగా ఉంది. You can choose a voice, change the pace, and try different languages. ముందుగా తెలుగులో ఒక వాక్యం విందాం, then we can switch to English.

temperature 0.8, top-k 30, top-p 1.0, maximum 3072 tokens. prompts and seeds.

benchmarks

to evaluate our specific use case, we need new benchmarks across different axes: expressiveness, consistency of non-verbal vocalizations, and basic WER/CER. there aren't any standardized benchmarks for multilingual indic TTS across these axes, so we present 3 benchmarks:

  • IndicEmo: evaluates expressiveness.
  • NoVA: evaluates adherence to non-verbal vocalization requests.
  • WER/CER: evaluates word and character error rates across 15 languages.

IndicEmo

github

we introduce IndicEmo to evaluate expressive delivery in code-switched speech. its 100 prompts combine two to five languages drawn from english, hindi, telugu, tamil, kannada, bengali, and punjabi, with equal coverage of happy, sad, angry, excited, and professional delivery.

three automated judges rate anonymized recordings on a five-point rubric covering tone, dynamics, phrasing, and sustained expression. results use the 98-prompt intersection with valid ratings from all judges, aggregating the median rating per recording into category means. rumik-oss 1 scores 2.92/5 overall and 3.03/5 on the four emotion categories.

systemoverall / 5emotions only / 5
gemini 3.1 flash tts preview4.584.66
rumik-oss 12.923.03
cartesia sonic preview2.712.51
cartesia sonic 3.52.322.18
elevenlabs eleven v32.162.03

the emotions-only sc

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys rumik-oss-1 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (rumik-oss-1 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="rumik-oss-1" -F file=@audio.mp3

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms