Model reference · open weights

wav2vec2-large-xlsr-53-th

wav2vec2-large-xlsr-53-th is an open-weight audio or speech model from airesearch, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio airesearch 1 variants 1.2M downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What wav2vec2-large-xlsr-53-th is

wav2vec2-large-xlsr-53-th Finetuning wav2vec2-large-xlsr-53 on Thai Common Voice 7.0 Read more on our blog We finetune wav2vec2-large-xlsr-53 based on Fine-tuning Wav2Vec2 for English ASR using Thai examples of Common Voice Corpus 7.0. The notebooks and scripts can be found in vistec-ai/wav2vec2-large-xlsr-53-th. The pretrained model and processor can be found at airesearch/wav2vec2-large-xlsr-53-th. robust-speech-event Add syllabletokenize, wordtokenize (PyThaiNLP) and deepcut tokenizers to eval.py from robust-speech-event Eval results on Common Voice 7 "test": Usage Datasets Common Voice Corpus 7.0](https://commonvoice.mozilla.org/en/datasets) contains 133 validated hours of Thai (255 total hours) at 5GB. We pre-tokenize with pythainlp.tokenize.wordtokenize. We preprocess the dataset using cleaning rules described in notebooks/cv-preprocess.ipynb by @tann9949. We then deduplicate and split as described in ekapolc/Thaicommonvoicesplit in order to 1) avoid data leakage due to random splits after cleaning in Common Voice Corpus 7.0 and 2) preserve the majority of the data for the training set. The dataset loading script is scripts/thcommonvoice70.py. You can use this scripts together with traincleand.tsv, validationcleaned.tsv and testcleaned.tsv to have the same splits as we do. The resulting dataset is as follows: Training We fintuned using the following configuration on a single V100 GPU and chose the checkpoint with the lowest validation loss. The finetuning script is scripts/wav2vec2finetune.py Evaluation We benchmark on the test set using WER with words tokenized by PyThaiNLP 2.3.1 and deepcut, and CER. We also measure performance when spell correction using TNC ngrams is applied. Evaluation codes can be found in notebooks/wav2vec2finetuningtutorial.ipynb. Benchmark is performed on test-unique split. ※ APIs are not finetuned with Common Voice 7.0 data LICENSE cc-by-sa 4.0 Ackowledgements model training and validation notebooks/scripts @cstorm125 dataset cleaning scripts @tann9949 dataset splits @ekapolc and @14mss running the training @mrpeerat spell correction @wannaphong

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makerairesearch
TypeAudio & music
Variants1
Runs withtransformers
Released2022-03-02
Popularity1.2M downloads / month
Likes28
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
wav2vec2-large-xlsr-53-thBF16Weights ↗

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
Automatic Speech RecognitionCommon Voice 7Test WER0.952
Automatic Speech RecognitionCommon Voice 7Test SER1.235
Automatic Speech RecognitionCommon Voice 7Test CER0.162

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys wav2vec2-large-xlsr-53-th for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (wav2vec2-large-xlsr-53-th below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="wav2vec2-large-xlsr-53-th" -F file=@audio.mp3

Details

Languages, data & research

Languages

th

Trained / evaluated on

common_voice

Tags

transformers pytorch wav2vec2 automatic-speech-recognition audio hf-asr-leaderboard robust-speech-event speech xlsr-fine-tuning th dataset:common_voice doi:10.57967/hf/0404 model-index endpoints_compatible

Licence

Open weights

Open weights under cc-by-sa-4.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want wav2vec2-large-xlsr-53-th on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms