Model reference · open weights
distill-neucodec is an open-weight audio or speech model from aoiandroid. distill-neucodec (BF16) weighs 1.0 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | aoiandroid |
|---|---|
| Type | Audio & music |
| Task | Audio→audio |
| Released | 2026-05-02 |
| Popularity | 532 downloads / month |
| Weights | 1.0 GB (distill-neucodec (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 1.0 GB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
Distill-NeuCodec is a version of NeuCodec with a compatible, distilled encoder.
The distilled encoder is 10x smaller in parameter count and uses ~7.5x less MACs at inference time.
The distilled model makes the following adjustments to the model:
Our work is largely based on extending the work of X-Codec2.0 and SQCodec.
Use the code below to get started with the model.
To install from pypi in a dedicated environment, using Python 3.10 or above:
conda create -n neucodec python=3.10
conda activate neucodec
pip install neucodec
Then, to use in python:
import librosa
import torch
import torchaudio
from torchaudio import transforms as T
from neucodec import DistillNeuCodec
model = DistillNeuCodec.from_pretrained("neuphonic/distill-neucodec")
model.eval().cuda()
y, sr = torchaudio.load(librosa.ex("libri1"))
if sr != 16_000:
y = T.Resample(sr, 16_000)(y)[None, ...] # (B, 1, T_16)
with torch.no_grad():
fsq_codes = model.encode_code(y)
# fsq_codes = model.encode_code(librosa.ex("libri1")) # or directly pass your filepath!
print(f"Codes shape: {fsq_codes.shape}")
recon = model.decode_code(fsq_codes).cpu() # (B, 1, T_24)
torchaudio.save("reconstructed.wav", recon[0, :, :], 24_000)
The model was trained using the same data as the full model, with an additional distillation loss (MSE between distilled and original encoder ouputs).
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works