Model reference · open weights

wav2vec2-xls-r-ca-lm

Available as managed deployment Audio PereLluis13 · community Speech→text 1 variants 3k dl/mo

wav2vec2-xls-r-ca-lm is an open-weight audio or speech model from PereLluis13. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byPereLluis13
TypeAudio & music
TaskSpeech→text
Runs withtransformers
Released2022-03-02
Popularity3k downloads / month
LicenceOpen weights

About

What wav2vec2-xls-r-ca-lm is

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on the MOZILLA-FOUNDATION/COMMON_VOICE_8_0 - CA, the tv3_parla and parlament_parla datasets.

Read the full model card

Model description

Please check the original facebook/wav2vec2-xls-r-1b Model card. This is just a finetuned version of that model.

Intended uses & limitations

As any model trained on crowdsourced data, this model can show the biases and particularities of the data and model used to train this model. Moreover, since this is a speech recognition model, it may underperform for some lower-resourced dialects for the catalan language.

Training and evaluation data

Training procedure

The data is preprocessed to remove characters not on the catalan alphabet. Moreover, numbers are verbalized using code provided by @ccoreilly, which can be found on the text/ folder or here.

Training results

Check the Tensorboard tab to check the training profile and evaluation results along training. The model was evaluated on the test splits for each of the datasets used during training.

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • seed: 42
  • gradient_accumulation_steps: 8
  • total_train_batch_size: 64
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 2000
  • num_epochs: 10.0
  • mixed_precision_training: Native AMP

Framework versions

  • Transformers 4.17.0.dev0
  • Pytorch 1.10.2+cu102
  • Datasets 1.18.3
  • Tokenizers 0.11.0

Thanks

Want to thank both @ccoreilly and @gullabi who have contributed with their own resources and knowledge into making this model possible.

From the published model card. Full card on the HuggingFace links in the sidebar.

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
Speech Recognitionmozilla-foundation/common_voice_8_0 caTest WER6.072
Speech Recognitionmozilla-foundation/common_voice_8_0 caTest CER1.918
Speech Recognitionprojecte-aina/parlament_parla caTest WER5.140
Speech Recognitionprojecte-aina/parlament_parla caTest CER2.016
Speech Recognitioncollectivat/tv3_parla caTest WER11.208
Speech Recognitioncollectivat/tv3_parla caTest CER7.321
Speech RecognitionRobust Speech Event - Catalan Dev DataTest WER22.870
Speech RecognitionRobust Speech Event - Catalan Dev DataTest CER13.590
Automatic Speech RecognitionRobust Speech Event - Test DataTest WER15.410

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys wav2vec2-xls-r-ca-lm for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (wav2vec2-xls-r-ca-lm below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="wav2vec2-xls-r-ca-lm" -F file=@audio.mp3

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms