Model reference · open weights
wav2vec2-french-phonemizer is an open-weight audio or speech model from Cnam-LMSSC. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | Cnam-LMSSC |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Parameters (lead) | 94M |
| Runs with | transformers |
| Released | 2026-01-08 |
| Popularity | 6k downloads / month |
| Licence | Open weights |
About
This version fixes the absence of the voiced labial-palatal approximant ɥ. The original Cnam-LMSSC/wav2vec2-french-phonemizer never transcribes the pronunciation of the voiced labial–palatal approximant with /ɥ/ and instead incorrectly replaces it with the sequence /yi/.
This created issues for phonetic accuracy and linguistic analysis. Words like "nuit" or "huit" are pronounced with \ɥ\ and not with \yi
For Text-to-Speech (TTS) task, substituting \ɥ\ with \yi\ results in unnatural prosody.
Fine-tuned facebook/wav2vec2-base-fr-voxpopuli-v2 for French speech-to-phoneme (without language model) using the train and validation splits of Common Voice v13.
When using this model, make sure that your speech input is sampled at 16kHz.
As this model is specifically trained for a speech-to-phoneme task, the output is sequence of IPA-encoded words, without punctuation. If you don't read the phonetic alphabet fluently, you can use this excellent IPA reader website to convert the transcript back to audio synthetic speech in order to check the quality of the phonetic transcription.
The model has been finetuned on Commonvoice-v13 (FR) for 14 epochs on a 1xADA_6000 GPU at Cnam/LMSSC using a ddp strategy and gradient-accumulation procedure (256 audios per update, corresponding roughly to 25 minutes of speech per update -> 2k updates per epoch)
Learning rate schedule : Double Tri-state schedule
The set of hyperparameters used for training are the same as those detailed in Annex B and Table 6 of wav2vec2 paper.
Just record your voice on the ⚡ Inference API on this webpage, and then click on "Compute", that's all !
The model can be used directly using the HuggingSound library:
import pandas as pd
from huggingsound import SpeechRecognitionModel
model = SpeechRecognitionModel("Cnam-LMSSC/wav2vec2-french-phonemizer-v2")
audio_paths = ["./test_relecture_texte.wav", "./10179_11051_000021.flac"]
# No need for the Audio files to be sampled at 16 kHz here,
# they are automatically resampled by Huggingsound
transcriptions = model.transcribe(audio_paths)
# (Optionnal) Display results in a table :
## transcriptions is list of dicts also containing timestamps and probabilities !
df = pd.DataFrame(transcriptions)
df['Audio file'] = pd.DataFrame(audio_paths)
df.set_index('Audio file', inplace=True)
df[['transcription']]
Output :
| Audio file | Phonetic transcription (IPA) |
|---|---|
| ./test_relecture_texte.wav | ʃapitʁ di də abɛse pəti kɔ̃t də ʒyl ləmɛtʁ ɑ̃ʁʒistʁe puʁ libʁivɔksɔʁɡ ibis dɑ̃ la bas kuʁ dœ̃ ʃato sə tʁuva paʁmi tut sɔʁt də volaj œ̃n ibis ʁɔz |
| ./10179_11051_000021.flac | kɛl dɔmaʒ kə sə nə swa pa dy sykʁ supiʁa se foʁaz ɑ̃ pasɑ̃ sa lɑ̃ɡ syʁ la vitʁ fɛ̃ dy ʃapitʁ kɛ̃z ɑ̃ʁʒistʁe paʁ sonjɛ̃ sɛt ɑ̃ʁʒistʁəmɑ̃ fɛ paʁti dy domɛn pyblik |
import torch
from transformers import AutoModelForCTC, Wav2Vec2Processor
from datasets import load_dataset
import soundfile as sf # Or Librosa if you prefer to ...
MODEL_ID = "Cnam-LMSSC/wav2vec2-french-phonemizer-v2"
model = AutoModelForCTC.from_pretrained(MODEL_ID)
processor = Wav2Vec2Processor.from_pretrained(MODEL_ID)
audio = sf.read('example.wav')
# Make sure you have a 16 kHz sampled audio file, or resample it !
inputs = processor(np.array(audio[0]),sampling_rate=16_000., return_tensors="pt")
with torch.no_grad():
logits = model(**inputs).logits
predicted_ids = torch.argmax(logits,dim = -1)
transcription = processor.batch_decode(predicted_ids)
print("Phonetic transcription : ", transcription)
Output :
'ʒə syi tʁɛ kɔ̃tɑ̃ də vu pʁezɑ̃te notʁ solysjɔ̃ puʁ fonomize dez odjo fasilmɑ̃ sa fɔ̃ksjɔn kɑ̃ mɛm tʁɛ bjɛ̃'
In the table below, we report the Phoneme Error Rate (PER) of the model on both Common Voice and Multilingual Librispeech (using the French configs for both datasets of course), when finetuned on Common Voice train set only :
| Model | Test Set | PER |
|---|---|---|
| Cnam-LMSSC/wav2vec2-french-phonemizer-v2 | Common Voice v13 (French) | 4.75% |
| Cnam-LMSSC/wav2vec2-french-phonemizer-v2 | Multilingual Librispeech (French) | 5.97% |
If you use this finetuned model for any publication, please use this to cite our work :
@misc {lmssc-wav2vec2-base-phonemizer-french-v2_2026,
author = { Olivier, Malo },
title = { wav2vec2-french-phonemizer-v2 (Revision 8e71385) },
year = 2026,
url = { https://huggingface.co/Cnam-LMSSC/wav2vec2-french-phonemizer-v2 },
doi = { 10.57967/hf/7468 },
publisher = { Hugging Face }
}
From the published model card. Full card on the HuggingFace links in the sidebar.
Benchmarks
As published on the model card — the maker's own numbers, not measured by AxForge.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| Speech Recognition | Common Voice v13 | Test PER on Common Voice FR 13.0 | Trained | 4.750 |
| Speech Recognition | Common Voice v13 | Test PER on Multilingual Librispeech FR | Trained | 5.970 |
| Speech Recognition | Common Voice v13 | Val PER on Common Voice FR 13.0 | Trained | 3.510 |
| Speech Recognition | Common Voice v13 | Val PER on Multilingual Librispeech FR | Trained | 5.070 |
Using it via the API
Once AxForge deploys wav2vec2-french-phonemizer for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (wav2vec2-french-phonemizer below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \ -H "Authorization: Bearer $AXFORGE_API_KEY" \ -F model="wav2vec2-french-phonemizer" -F file=@audio.mp3
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.