Model reference · open weights
wav2vec2-xlarge-fi-150k-finetuned is an open-weight audio or speech model from GetmanY1. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | GetmanY1 |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Parameters (lead) | 963M |
| Runs with | transformers |
| Based on | GetmanY1/wav2vec2-xlarge-fi-150k |
| Released | 2024-09-13 |
| Popularity | 524 downloads / month |
| Licence | Open weights |
About
GetmanY1/wav2vec2-xlarge-fi-150k fine-tuned on 4600 hours of Finnish speech on 16kHz sampled speech audio:
When using the model make sure that your speech input is also sampled at 16Khz.
The Finnish Wav2Vec2 X-Large has the same architecture and uses the same training objective as the multilingual one described in paper.
GetmanY1/wav2vec2-xlarge-fi-150k is a large-scale, 1-billion parameter monolingual model pre-trained on 158k hours of unlabeled Finnish speech, including KAVI radio and television archive materials, Lahjoita puhetta (Donate Speech), Finnish Parliament, Finnish VoxPopuli.
You can read more about the pre-trained model from this paper. The training scripts are available on GitHub.
You can use this model for Finnish ASR (speech-to-text).
To transcribe audio files the model can be used as a standalone acoustic model as follows:
from transformers import Wav2Vec2Processor, Wav2Vec2ForCTC
from datasets import load_dataset
import torch
# load model and processor
processor = Wav2Vec2Processor.from_pretrained("GetmanY1/wav2vec2-xlarge-fi-150k-finetuned")
model = Wav2Vec2ForCTC.from_pretrained("GetmanY1/wav2vec2-xlarge-fi-150k-finetuned")
# load dummy dataset and read soundfiles
ds = load_dataset("mozilla-foundation/common_voice_16_1", "fi", split='test')
# tokenize
input_values = processor(ds[0]["audio"]["array"], return_tensors="pt", padding="longest").input_values # Batch size 1
# retrieve logits
logits = model(input_values).logits
# take argmax and decode
predicted_ids = torch.argmax(logits, dim=-1)
transcription = processor.batch_decode(predicted_ids)
If you use our models or scripts, please cite our article as:
@inproceedings{getman25_interspeech,
title = {{Is your model big enough? Training and interpreting large-scale monolingual speech foundation models}},
author = {{Yaroslav Getman and Tamás Grósz and Tommi Lehtonen and Mikko Kurimo}},
year = {{2025}},
booktitle = {{Interspeech 2025}},
pages = {{231--235}},
doi = {{10.21437/Interspeech.2025-46}},
issn = {{2958-1796}},
}
Feel free to contact us for more details 🤗
From the published model card. Full card on the HuggingFace links in the sidebar.
Benchmarks
As published on the model card — the maker's own numbers, not measured by AxForge.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| Automatic Speech Recognition | Lahjoita puhetta (Donate Speech) | Dev WER | 14.980 |
| Automatic Speech Recognition | Lahjoita puhetta (Donate Speech) | Dev CER | 4.130 |
| Automatic Speech Recognition | Lahjoita puhetta (Donate Speech) | Test WER | 16.370 |
| Automatic Speech Recognition | Lahjoita puhetta (Donate Speech) | Test CER | 5.030 |
| Automatic Speech Recognition | Finnish Parliament | Dev16 WER | 10.910 |
| Automatic Speech Recognition | Finnish Parliament | Dev16 CER | 4.850 |
| Automatic Speech Recognition | Finnish Parliament | Test16 WER | 7.810 |
| Automatic Speech Recognition | Finnish Parliament | Test16 CER | 3.480 |
| Automatic Speech Recognition | Finnish Parliament | Test20 WER | 6.430 |
| Automatic Speech Recognition | Finnish Parliament | Test20 CER | 2.090 |
| Automatic Speech Recognition | Common Voice 16.1 | Dev WER | 6.650 |
| Automatic Speech Recognition | Common Voice 16.1 | Dev CER | 1.150 |
| Automatic Speech Recognition | Common Voice 16.1 | Test WER | 5.420 |
| Automatic Speech Recognition | Common Voice 16.1 | Test CER | 0.960 |
| Automatic Speech Recognition | FLEURS | Dev WER | 8.670 |
| Automatic Speech Recognition | FLEURS | Dev CER | 5.180 |
| Automatic Speech Recognition | FLEURS | Test WER | 9.960 |
| Automatic Speech Recognition | FLEURS | Test CER | 5.740 |
Using it via the API
Once AxForge deploys wav2vec2-xlarge-fi-150k-finetuned for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (wav2vec2-xlarge-fi-150k-finetuned below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \ -H "Authorization: Bearer $AXFORGE_API_KEY" \ -F model="wav2vec2-xlarge-fi-150k-finetuned" -F file=@audio.mp3
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.