Model reference · open weights

wav2vec2-xls-r-mixed

wav2vec2-xls-r-mixed is an open-weight audio or speech model from mesolitica, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Licence fee required Audio mesolitica 1 variants 1.1M downloads/mo
Request a licence + hosting quote All served models Not on the shared API today — deployed on request.

About

What wav2vec2-xls-r-mixed is

probably proofread and complete it, then remove this comment. -- wav2vec2-xls-r-300m-mixed Finetuned https://huggingface.co/facebook/wav2vec2-xls-r-300m on https://github.com/huseinzol05/malaya-speech/tree/master/data/mixed-stt This model was finetuned on 3 languages, 1. Malay 2. Singlish 3. Mandarin This model trained on a single RTX 3090 Ti 24GB VRAM, provided by https://mesolitica.com/. Evaluation set Evaluation set from https://github.com/huseinzol05/malaya-speech/tree/master/pretrained-model/prepare-stt with sizes, It achieves the following results on the evaluation set based on evaluate-gpu.ipynb: Mixed evaluation, Malay evaluation, Singlish evaluation, Mandarin evaluation, Language model from https://huggingface.co/huseinzol05/language-model-bahasa-manglish-combined

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makermesolitica
TypeAudio & music
Variants1
Runs withtransformers
Released2022-06-01
Popularity1.1M downloads / month
Likes5
LicenceCommercial licence needed

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
wav2vec2-xls-r-300m-mixedBF16Weights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys wav2vec2-xls-r-mixed for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (wav2vec2-xls-r-mixed below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="wav2vec2-xls-r-mixed" -F file=@audio.mp3

Details

Languages, data & research

Tags

transformers pytorch tf wav2vec2 automatic-speech-recognition generated_from_keras_callback endpoints_compatible deploy:azure

Licence

Commercial licence needed

The weights are open but its licence needs a commercial agreement for business use. AxForge can arrange that licence and host the model for you — you pay AxForge, we settle with the model’s maker. Ask us for a quote. Read the licence ↗

Sources

Weights & code

Want wav2vec2-xls-r-mixed on EU-owned hardware?

Request a licence + hosting quote See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms