Model reference · open weights
wav2vec2-xls-r-cv8-turkish is an open-weight audio or speech model from mpoyraz. wav2vec2-xls-r-300m-cv8-turkish (BF16) weighs 1.3 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | mpoyraz |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Runs with | transformers |
| Released | 2022-03-02 |
| Popularity | 605 downloads / month |
| Weights | 1.3 GB (wav2vec2-xls-r-300m-cv8-turkish (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 1.3 GB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
This ASR model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on Turkish language.
The following datasets were used for finetuning:
validated split except test split was used for training.To support the datasets above, custom pre-processing and loading steps was performed and wav2vec2-turkish repo was used for that purpose.
The following hypermaters were used for finetuning:
N-gram language model is trained on a Turkish Wikipedia articles using KenLM and ngram-lm-wiki repo was used to generate arpa LM and convert it into binary format.
Please install unicode_tr package before running evaluation. It is used for Turkish text processing.
mozilla-foundation/common_voice_8_0 with split testpython eval.py --model_id mpoyraz/wav2vec2-xls-r-300m-cv8-turkish --dataset mozilla-foundation/common_voice_8_0 --config tr --split test
speech-recognition-community-v2/dev_datapython eval.py --model_id mpoyraz/wav2vec2-xls-r-300m-cv8-turkish --dataset speech-recognition-community-v2/dev_data --config tr --split validation --chunk_length_s 5.0 --stride_length_s 1.0
| Dataset | WER | CER |
|---|---|---|
| Common Voice 8 TR test split | 10.61 | 2.67 |
| Speech Recognition Community dev data | 36.46 | 12.38 |
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
Benchmarks
As published on the model card — the maker's own numbers, not measured by AxForge.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| Automatic Speech Recognition | Common Voice 8 | Test WER | 10.610 |
| Automatic Speech Recognition | Common Voice 8 | Test CER | 2.670 |
| Automatic Speech Recognition | Robust Speech Event - Dev Data | Test WER | 36.460 |
| Automatic Speech Recognition | Robust Speech Event - Dev Data | Test CER | 12.380 |
| Automatic Speech Recognition | Robust Speech Event - Test Data | Test WER | 40.910 |