Model reference · open weights
distill-whisper-th-small is an open-weight audio or speech model from biodatlab. distill-whisper-th-small (FP16) weighs 412 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | biodatlab |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Parameters (lead) | 206M |
| Runs with | transformers |
| Released | 2024-01-16 |
| Popularity | 1k downloads / month |
| Weights | 412 MB (distill-whisper-th-small (FP16), file size) |
| Licence | Open weights |
What it runs on
Weights 412 MB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
This is a distilled Automatic Speech Recognition (ASR) model, based on the Whisper architecture. It has been specifically tailored for Thai language speech recognition. The model features 4 decoder layers (vs 12 in teacher model) and has been distilled from a larger teacher model, focusing on enhancing performance and efficiency.
This shows an improvement in Word Error Rate (WER), indicating enhanced accuracy in speech recognition tasks for the Thai language.
This model is intended for use in applications requiring Thai language speech recognition.
This model was developed using resources and datasets provided by the speech and language technology community. Special thanks to the teams behind Common Voice, Gowajee, SLSCU, and the Thai Elderly Speech Corpus for their valuable datasets.
Cite using Bibtex:
@misc {thonburian_whisper_med,
author = { Atirut Boribalburephan, Zaw Htet Aung, Knot Pipatsrisawat, Titipat Achakulvisut },
title = { Thonburian Whisper: A fine-tuned Whisper model for Thai automatic speech recognition },
year = 2022,
url = { https://huggingface.co/biodatlab/distil-whisper-th-small },
doi = { 10.57967/hf/0226 },
publisher = { Hugging Face }
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.