Model reference · open weights
nepali-voice-engine is an open-weight audio or speech model from ujjwal5454. nepali-voice-engine-v4 (FP32) weighs 1.9 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | ujjwal5454 |
|---|---|
| Type | Audio & music |
| Task | Text→speech |
| Parameters (lead) | 938M |
| Runs with | transformers |
| Released | 2026-09-26 |
| Popularity | 661 downloads / month |
| Weights | 1.9 GB (nepali-voice-engine-v4 (FP32), file size) |
| Licence | Licence not stated |
What it runs on
Weights 1.9 GB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
A self-contained Transformers remote-code package for the v4 checkpoint, with Aakriti, Sangeeta, Prabal and Baje persona presets. The release contains the speech decoder, both tokenizers and embedded text-encoder/audio-codec configurations. T5 and DAC run through standard Transformers; no separate TTS framework or codec package is imported or downloaded.
pip install torch 'transformers==4.46.1'
This version is pinned because the decoder uses Transformers' internal generation and cache interfaces. Transformers 5.x is not supported by this release. PyTorch and Transformers' normal transitive dependencies are required; “self-contained” means no additional model/framework packages or external model repositories.
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained(
"ujjwal5454/nepali-voice-engine-v4", trust_remote_code=True
)
# Optional on a CUDA machine:
# model = model.to("cuda")
audio = model.generate_speech(prompt="नमस्ते सबैलाई!", persona="aakriti")
sample_rate = model.sampling_rate # 44100
audio is a one-dimensional CPU float32 PyTorch tensor. Personas are aakriti,
sangeeta, prabal, and baje (case insensitive). Generation overrides are
accepted, for example max_new_tokens=1200. The default 800-token budget can
truncate long speech. Presets condition the model with text descriptions; they
are not hard speaker-ID guarantees. No postprocessing filter changes the waveform.
AutoTokenizer.from_pretrained(...) loads the Nepali prompt tokenizer; the engine
loads the bundled description_tokenizer/ automatically from the same revision.
Once the complete release is downloaded, local loading works offline:
model = AutoModel.from_pretrained(
"./complete-release", trust_remote_code=True, local_files_only=True
)
config.json is the active Hub config; sanitized_config.json is an identical
review copy. Both use NepaliVoiceConfig, model type nepali_voice, and the
NepaliVoiceEngine architecture. The weights retain their original tensor names.
The two Python modules load submodels from embedded configs, without repository
lookups for T5 or DAC.
This directory is an overlay for the existing Hub repository. Its existing
model.safetensors must remain present in the published repository. To turn the
overlay into a complete local checkpoint, stage the original weights:
python tools/stage_weights.py /path/to/v4-snapshot
python tools/validate_release.py --load-weights --synthesize
To check the configuration, tokenizers, import boundaries, and all original weight shapes without allocating the full checkpoint:
python tools/validate_release.py
python tools/test_tiny_inference.py
The tiny test exercises autoregressive decoding and DAC synthesis with random
weights, plus save/reload in offline mode. It does not measure trained voice quality.
See PACKAGING_PLAN.md and VALIDATION.md for scope and validation status.
The custom decoder is an Apache-2.0 licensed derivative of Hugging Face's
Parler-TTS implementation.
The trained checkpoint derives from
Indic Parler-TTS.
Runtime independence does not erase provenance; see NOTICE and LICENSE.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.