Model reference · open weights

wav2vec2-large-960h-lv60-self

Audio facebook Speech→text 1 build Open weights 104k dl/mo

wav2vec2-large-960h-lv60-self is an open-weight audio or speech model from Meta. wav2vec2-large-960h-lv60-self (BF16) weighs 1.3 GB; the smallest configuration that runs it is RTX 3060 12 GB.

wav2vec2-large-960h-lv60-self is an automatic speech recognition model developed by Meta and published as facebook/wav2vec2-large-960h-lv60-self. It is designed for English speech and was pretrained and fine-tuned on 960 hours of Libri-Light and Librispeech audio sampled at 16kHz. The model utilizes a self-training objective and is distributed under the apache-2.0 license.

Summary of the facebook/wav2vec2-large-960h-lv60-self model card, 2026-10-01

What it is

Released byMeta
Published underfacebook
TypeAudio & music
TaskSpeech→text
Runs withtransformers
Released2022-03-02
Popularity104k downloads / month
Weights1.3 GB (wav2vec2-large-960h-lv60-self (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for wav2vec2-large-960h-lv60-self (BF16)

Weights 1.3 GB (file size) · overhead about 1.6 GB.

CardOne streamCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What Meta says about wav2vec2-large-960h-lv60-self

Read the model card

Facebook's Wav2Vec2

The large model pretrained and fine-tuned on 960 hours of Libri-Light and Librispeech on 16kHz sampled speech audio. Model was trained with Self-Training objective. When using the model make sure that your speech input is also sampled at 16Khz.

Paper

Authors: Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, Michael Auli

Abstract

We show for the first time that learning powerful representations from speech audio alone followed by fine-tuning on transcribed speech can outperform the best semi-supervised methods while being conceptually simpler. wav2vec 2.0 masks the speech input in the latent space and solves a contrastive task defined over a quantization of the latent representations which are jointly learned. Experiments using all labeled data of Librispeech achieve 1.8/3.3 WER on the clean/other test sets. When lowering the amount of labeled data to one hour, wav2vec 2.0 outperforms the previous state of the art on the 100 hour subset while using 100 times less labeled data. Using just ten minutes of labeled data and pre-training on 53k hours of unlabeled data still achieves 4.8/8.2 WER. This demonstrates the feasibility of speech recognition with limited amounts of labeled data.

The original model can be found under https://github.com/pytorch/fairseq/tree/master/examples/wav2vec#wav2vec-20.

Usage

To transcribe audio files the model can be used as a standalone acoustic model as follows:

 from transformers import Wav2Vec2Processor, Wav2Vec2ForCTC
 from datasets import load_dataset
 import torch

 # load model and processor
 processor = Wav2Vec2Processor.from_pretrained("facebook/wav2vec2-large-960h-lv60-self")
 model = Wav2Vec2ForCTC.from_pretrained("facebook/wav2vec2-large-960h-lv60-self")

 # load dummy dataset and read soundfiles
 ds = load_dataset("patrickvonplaten/librispeech_asr_dummy", "clean", split="validation")

 # tokenize
 input_values = processor(ds[0]["audio"]["array"], return_tensors="pt", padding="longest").input_values

 # retrieve logits
 logits = model(input_values).logits

 # take argmax and decode
 predicted_ids = torch.argmax(logits, dim=-1)
 transcription = processor.batch_decode(predicted_ids)

Evaluation

This code snippet shows how to evaluate facebook/wav2vec2-large-960h-lv60-self on LibriSpeech's "clean" and "other" test data.

from datasets import load_dataset
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
import torch
from jiwer import wer

librispeech_eval = load_dataset("librispeech_asr", "clean", split="test")

model = Wav2Vec2ForCTC.from_pretrained("facebook/wav2vec2-large-960h-lv60-self").to("cuda")
processor = Wav2Vec2Processor.from_pretrained("facebook/wav2vec2-large-960h-lv60-self")

def map_to_pred(batch):
    inputs = processor(batch["audio"]["array"], return_tensors="pt", padding="longest")
    input_values = inputs.input_values.to("cuda")
    attention_mask = inputs.attention_mask.to("cuda")

    with torch.no_grad():
        logits = model(input_values, attention_mask=attention_mask).logits

    predicted_ids = torch.argmax(logits, dim=-1)
    transcription = processor.batch_decode(predicted_ids)
    batch["transcription"] = transcription
    return batch

result = librispeech_eval.map(map_to_pred, remove_columns=["audio"])

print("WER:", wer(result["text"], result["transcription"]))

Result (WER):

"clean""other"
1.93.9

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
Automatic Speech RecognitionLibriSpeech (clean)Test WER1.900
Automatic Speech RecognitionLibriSpeech (other)Test WER3.900
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms