Model reference · open weights
Whisperv3-tunisian-codeswitch is an open-weight audio or speech model from oddadmix. Whisperv3-tunisian-codeswitch (FP32) weighs 3.1 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | oddadmix |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Parameters (lead) | 1.5B |
| Based on | oddadmix/whisper-large-v3-tunisian-codeswitch-asr-v2 |
| Released | 2026-07-28 |
| Popularity | 673 downloads / month |
| Weights | 3.1 GB (Whisperv3-tunisian-codeswitch (FP32), file size) |
| Licence | Licence not stated |
What it runs on
Weights 3.1 GB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
Whisper-large-v3 full fine-tune for Tunisian Arabic ↔ French/English code-switched ASR (NADI 2026 shared task, subtask 1.3).
tun-asr-aug-v4 on FARUKxAUTO/tunisian-asr-cleaned (46K, dense
Tunisian↔French code-switch) + NADI TEDx train replay, 2 epochs, --spec_augment,
paged_adamw_8bit, lr 5e-6 (a single longer cosine schedule — 2 epochs is the sweet spot).language="ar", task="transcribe". Submit raw (scorer applies clean_transcription).from transformers import WhisperForConditionalGeneration, WhisperProcessor
import torch
m = WhisperForConditionalGeneration.from_pretrained("oddadmix/nadi2026-subtask1.3-tunisian-codeswitch-asr-faruk-v7", torch_dtype=torch.bfloat16).cuda().eval()
p = WhisperProcessor.from_pretrained("oddadmix/nadi2026-subtask1.3-tunisian-codeswitch-asr-faruk-v7")
# feats = p(audio, sampling_rate=16000, return_tensors="pt").input_features.cuda().to(torch.bfloat16)
# ids = m.generate(feats, language="ar", task="transcribe", max_new_tokens=256)
# print(p.batch_decode(ids, skip_special_tokens=True))
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.