Model reference · open weights
MMS-TTS-THAI-MALEV2 is an open-weight audio or speech model from VIZINTZOR. MMS-TTS-THAI-MALEV2 (FP32) weighs 166 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | VIZINTZOR |
|---|---|
| Type | Audio & music |
| Task | Text→speech |
| Parameters (lead) | 83M |
| Runs with | transformers |
| Released | 2025-01-25 |
| Popularity | 506 downloads / month |
| Weights | 166 MB (MMS-TTS-THAI-MALEV2 (FP32), file size) |
| Licence | Licence not stated |
What it runs on
Weights 166 MB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
โมเดลนี้ใช้ เสียงที่บันทึกจาก Play.ht : https://play.ht/ เพื่อนำมา finetune model.
Finetune โมเดลโค้ด GitHub : https://github.com/VYNCX/finetune-local-vits
) เทรนโมเดลเสียงด้วยตัวเองบน Google Colab
ใช้งาน บน local คอมพิวเตอร์ https://github.com/VYNCX/VachanaTTS
การใช้งาน :
import torch
from transformers import VitsTokenizer, VitsModel, set_seed
import scipy
tokenizer = VitsTokenizer.from_pretrained("VIZINTZOR/MMS-TTS-THAI-MALEV2",cache_dir="./mms")
model = VitsModel.from_pretrained("VIZINTZOR/MMS-TTS-THAI-MALEV2",cache_dir="./mms")
inputs = tokenizer(text="สวัสดีครับ นี่คือเสียงพูดภาษาไทย", return_tensors="pt")
set_seed(456) # make deterministic
with torch.no_grad():
outputs = model(**inputs)
waveform = outputs.waveform[0]
# Convert PyTorch tensor to NumPy array
waveform_array = waveform.numpy()
scipy.io.wavfile.write("techno_output.wav", rate=model.config.sampling_rate, data=waveform_array)
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.