Model reference · open weights

nepali-voice-engine

NEW · this week Audio ujjwal5454 · community Text→speech 1 build Licence not stated 661 dl/mo

nepali-voice-engine is an open-weight audio or speech model from ujjwal5454. nepali-voice-engine-v4 (FP32) weighs 1.9 GB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byujjwal5454
TypeAudio & music
TaskText→speech
Parameters (lead)938M
Runs withtransformers
Released2026-09-26
Popularity661 downloads / month
Weights1.9 GB (nepali-voice-engine-v4 (FP32), file size)
LicenceLicence not stated

What it runs on

Memory and cards for nepali-voice-engine-v4 (FP32)

Weights 1.9 GB (file size) · overhead about 1.6 GB.

CardOne streamCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What ujjwal5454 says about nepali-voice-engine

A self-contained Transformers remote-code package for the v4 checkpoint, with Aakriti, Sangeeta, Prabal and Baje persona presets. The release contains the speech decoder, both tokenizers and embedded text-encoder/audio-codec configurations. T5 and DAC run through standard Transformers; no separate TTS framework or codec package is imported or downloaded.

Read the full model card

Install and use

pip install torch 'transformers==4.46.1'

This version is pinned because the decoder uses Transformers' internal generation and cache interfaces. Transformers 5.x is not supported by this release. PyTorch and Transformers' normal transitive dependencies are required; “self-contained” means no additional model/framework packages or external model repositories.

from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained(
    "ujjwal5454/nepali-voice-engine-v4", trust_remote_code=True
)
# Optional on a CUDA machine:
# model = model.to("cuda")
audio = model.generate_speech(prompt="नमस्ते सबैलाई!", persona="aakriti")
sample_rate = model.sampling_rate  # 44100

audio is a one-dimensional CPU float32 PyTorch tensor. Personas are aakriti, sangeeta, prabal, and baje (case insensitive). Generation overrides are accepted, for example max_new_tokens=1200. The default 800-token budget can truncate long speech. Presets condition the model with text descriptions; they are not hard speaker-ID guarantees. No postprocessing filter changes the waveform.

AutoTokenizer.from_pretrained(...) loads the Nepali prompt tokenizer; the engine loads the bundled description_tokenizer/ automatically from the same revision. Once the complete release is downloaded, local loading works offline:

model = AutoModel.from_pretrained(
    "./complete-release", trust_remote_code=True, local_files_only=True
)

Release contents and verification

config.json is the active Hub config; sanitized_config.json is an identical review copy. Both use NepaliVoiceConfig, model type nepali_voice, and the NepaliVoiceEngine architecture. The weights retain their original tensor names. The two Python modules load submodels from embedded configs, without repository lookups for T5 or DAC.

This directory is an overlay for the existing Hub repository. Its existing model.safetensors must remain present in the published repository. To turn the overlay into a complete local checkpoint, stage the original weights:

python tools/stage_weights.py /path/to/v4-snapshot
python tools/validate_release.py --load-weights --synthesize

To check the configuration, tokenizers, import boundaries, and all original weight shapes without allocating the full checkpoint:

python tools/validate_release.py
python tools/test_tiny_inference.py

The tiny test exercises autoregressive decoding and DAC synthesis with random weights, plus save/reload in offline mode. It does not measure trained voice quality. See PACKAGING_PLAN.md and VALIDATION.md for scope and validation status.

The custom decoder is an Apache-2.0 licensed derivative of Hugging Face's Parler-TTS implementation. The trained checkpoint derives from Indic Parler-TTS. Runtime independence does not erase provenance; see NOTICE and LICENSE.

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms