Model reference · open weights
MioCodec-25Hz-24kHz is an open-weight audio or speech model from Aratako. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Maker | Aratako |
|---|---|
| Type | Audio & music |
| Task | Audio→audio |
| Parameters (lead) | 131M |
| Released | 2026-02-03 |
| Popularity | 14k downloads / month |
| Licence | Open weights |
About
MioCodec-25Hz-24kHz is a lightweight and fast neural audio codec designed for efficient spoken language modeling. Based on the Kanade-Tokenizer implementation, this model features an integrated wave decoder (iSTFTHead) that directly synthesizes waveforms without requiring an external vocoder.
For higher audio fidelity at 44.1 kHz, see MioCodec-25Hz-44.1kHz.
MioCodec decomposes speech into two distinct components:
By disentangling these elements, MioCodec is ideal for Spoken Language Modeling.
| Model | Token Rate | Vocab Size | Bit Rate | Sample Rate | SSL Encoder | Vocoder | Parameters | Highlights |
|---|---|---|---|---|---|---|---|---|
| MioCodec-25Hz-24kHz | 25 Hz | 12,800 | 341 bps | 24 kHz | WavLM-base+ | - (iSTFTHead) | 132M | Lightweight, fast inference |
| MioCodec-25Hz-44.1kHz | 25 Hz | 12,800 | 341 bps | 44.1 kHz | WavLM-base+ | MioVocoder (Jointly Tuned) | 118M (w/o vocoder) | High-quality, high sample rate |
| kanade-25hz | 25 Hz | 12,800 | 341 bps | 24 kHz | WavLM-base+ | Vocos 24kHz | 118M (w/o vocoder) | Original 25Hz model |
| kanade-12.5hz | 12.5 Hz | 12,800 | 171 bps | 24 kHz | WavLM-base+ | Vocos 24kHz | 120M (w/o vocoder) | Original 12.5Hz model |
# Install via pip
pip install git+https://github.com/Aratako/MioCodec
# Or using uv
uv add git+https://github.com/Aratako/MioCodec
Basic usage for encoding and decoding audio:
from miocodec import MioCodecModel, load_audio
import soundfile as sf
# 1. Load model
model = MioCodecModel.from_pretrained("Aratako/MioCodec-25Hz-24kHz").eval().cuda()
# 2. Load audio
waveform = load_audio("input.wav", sample_rate=model.config.sample_rate).cuda()
# 3. Encode Audio
features = model.encode(waveform)
# 4. Decode to Waveform (directly, no vocoder needed)
resynth = model.decode(
content_token_indices=features.content_token_indices,
global_embedding=features.global_embedding,
)
# 5. Save
sf.write("output.wav", resynth.cpu().numpy(), model.config.sample_rate)
MioCodec allows you to swap speaker identities by combining the content tokens of a source with the global embedding of a reference.
source = load_audio("source_content.wav", sample_rate=model.config.sample_rate).cuda()
reference = load_audio("target_speaker.wav", sample_rate=model.config.sample_rate).cuda()
# Perform conversion
vc_wave = model.voice_conversion(source, reference)
sf.write("converted.wav", vc_wave.cpu().numpy(), model.config.sample_rate)
MioCodec-25Hz-24kHz was trained in two phases with an integrated wave decoder that directly synthesizes waveforms via iSTFT.
The model is trained to minimize both Multi-Resolution Mel-spectrogram loss and SSL feature reconstruction loss (using WavLM-base+). The wave decoder directly generates waveforms, and losses are computed on the reconstructed audio.
[32, 64, 128, 256, 512, 1024, 2048].Building upon Phase 1, adversarial training is introduced to improve perceptual quality. The training objectives include:
[32, 64, 128, 256, 512, 1024, 2048].[2, 3, 5, 7, 11, 17, 23].[118, 190, 310, 502, 814, 1314, 2128, 3444].The training datasets are listed below:
| Language | Approx. Hours | Dataset |
|---|---|---|
| Japanese | ~22,500h | Various public HF datasets |
| English | ~500h | Libriheavy-HQ |
| English | ~4,000h | MLS-Sidon |
| English | ~9,000h | HiFiTTS-2 |
| English | ~27,000h | Emilia-YODAS |
| German | ~1,950h | MLS-Sidon |
| German | ~5,600h | Emilia-YODAS |
| Dutch | ~1,550h | MLS-Sidon |
| French | ~1,050h | MLS-Sidon |
| French | ~7,400h | Emilia-YODAS |
| Spanish | ~900h | MLS-Sidon |
| Italian | ~240h | MLS-Sidon |
| Portuguese | ~160h | [MLS-Sidon](ht |
From the published model card. Full card on the HuggingFace links in the sidebar.
How it works
Using it via the API
Once AxForge deploys miocodec-25hz-24khz for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (miocodec-25hz-24khz below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \ -H "Authorization: Bearer $AXFORGE_API_KEY" \ -F model="miocodec-25hz-24khz" -F file=@audio.mp3
Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.