Model reference · open weights
LongCat-AudioDiT is an open-weight audio or speech model from drbaph. LongCat-AudioDiT-3.5B-bf16 (BF16) weighs 7.7 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | drbaph |
|---|---|
| Type | Audio & music |
| Task | Text→speech |
| Parameters (lead) | 3.8B |
| Runs with | transformers |
| Released | 2026-03-30 |
| Popularity | 5k downloads / month |
| Weights | 7.7 GB (LongCat-AudioDiT-3.5B-bf16 (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 7.7 GB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
BF16 quantized version of meituan-longcat/LongCat-AudioDiT-3.5B.
Original Model | Paper | GitHub (Original) | ComfyUI Node
This is a BF16 conversion of LongCat-AudioDiT-3.5B — a state-of-the-art diffusion-based zero-shot TTS model by Meituan that operates directly in the waveform latent space. Converting from FP32 to BF16 halves the on-disk size and VRAM usage with negligible quality loss, making it the recommended variant for most users.
| Original (3.5B FP32) | This (3.5B BF16) | |
|---|---|---|
| Weight dtype | float32 | bfloat16 |
| Activation dtype | float32 | bfloat16 |
| File size | ~14 GB | ~7 GB |
| VRAM (inference) | ~20 GB | ~12 GB |
| Quality | Reference | Virtually identical |
| Extra dependencies | none | none |
All model weights — DiT transformer backbone, Wav-VAE, and text encoder — are converted from float32 to bfloat16. BF16 preserves the same dynamic range as FP32 (8 exponent bits) while halving memory usage, making it the lossless practical choice for inference on modern GPUs.
No post-training quantization, calibration data, or scale factors are required. The model is a direct dtype cast and is fully compatible with the original audiodit inference code.
The easiest way to use this model is with ComfyUI-LongCat-AudioDIT-TTS, which has native support for this BF16 model with zero extra setup.
LongCat-AudioDiT) or manually: cd ComfyUI/custom_nodes
git clone https://github.com/Saganaki22/ComfyUI-LongCat-AudioDIT-TTS.git
The model auto-downloads on first use — select LongCat-AudioDiT-3.5B-bf16 from the model dropdown in any LongCat node.
Or download manually:
huggingface-cli download drbaph/LongCat-AudioDiT-3.5B-bf16 --local-dir ComfyUI/models/audiodit/LongCat-AudioDiT-3.5B-bf16
dtype: auto or bf16 — matches this model's native dtypeguidance_method: cfg for TTS, apg for voice cloningsteps: 16 (balanced), 32 (higher quality)keep_model_loaded: True for repeated useThis is the recommended variant for most users — best balance of quality, VRAM usage, and compatibility.
LongCat-AudioDiT is a non-autoregressive diffusion-based TTS model from Meituan that achieves state-of-the-art zero-shot voice cloning performance on the Seed benchmark. Unlike previous methods relying on mel-spectrograms, it operates directly in the waveform latent space using only a Wav-VAE and a DiT backbone.
The 3.5B variant achieves 0.818 SIM on Seed-ZH and 0.797 SIM on Seed-Hard, surpassing both open-source and closed-source competitors.
This model inherits the MIT License from meituan-longcat/LongCat-AudioDiT-3.5B.
The BF16 conversion was produced by drbaph and is released under the same license.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works