Model reference · open weights

wav2vec2-large-xlsr-53-polish

wav2vec2-large-xlsr-53-polish is an open-weight audio or speech model from jonatasgrosman, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio jonatasgrosman 1 variants 3.1M downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What wav2vec2-large-xlsr-53-polish is

Fine-tuned XLSR-53 large model for speech recognition in Polish Fine-tuned facebook/wav2vec2-large-xlsr-53 on Polish using the train and validation splits of Common Voice 6.1. When using this model, make sure that your speech input is sampled at 16kHz. This model has been fine-tuned thanks to the GPU credits generously given by the OVHcloud :) The script used for training can be found here: https://github.com/jonatasgrosman/wav2vec2-sprint Usage The model can be used directly (without a language model) as follows... Using the HuggingSound library: Writing your own inference script: Evaluation 1. To evaluate on mozilla-foundation/commonvoice60 with split test 2. To evaluate on speech-recognition-community-v2/devdata Citation If you want to cite this model you can use this:

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makerjonatasgrosman
TypeAudio & music
Variants1
Runs withtransformers
Released2022-03-02
Popularity3.1M downloads / month
Likes12
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
wav2vec2-large-xlsr-53-polishBF16Weights ↗

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
Automatic Speech RecognitionCommon Voice plTest WER14.21
Automatic Speech RecognitionCommon Voice plTest CER3.49
Automatic Speech RecognitionCommon Voice plTest WER (+LM)10.98
Automatic Speech RecognitionCommon Voice plTest CER (+LM)2.93
Automatic Speech RecognitionRobust Speech Event - Dev DataDev WER33.18
Automatic Speech RecognitionRobust Speech Event - Dev DataDev CER15.92
Automatic Speech RecognitionRobust Speech Event - Dev DataDev WER (+LM)29.31
Automatic Speech RecognitionRobust Speech Event - Dev DataDev CER (+LM)15.17

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys wav2vec2-large-xlsr-53-polish for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (wav2vec2-large-xlsr-53-polish below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="wav2vec2-large-xlsr-53-polish" -F file=@audio.mp3

Details

Languages, data & research

Languages

pl

Trained / evaluated on

common_voice mozilla-foundation/common_voice_6_0

Tags

transformers pytorch jax wav2vec2 automatic-speech-recognition audio hf-asr-leaderboard mozilla-foundation/common_voice_6_0 pl robust-speech-event speech xlsr-fine-tuning-week dataset:common_voice dataset:mozilla-foundation/common_voice_6_0

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want wav2vec2-large-xlsr-53-polish on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms