Model reference · open weights

parakeet-tdt

parakeet-tdt is an open-weight audio or speech model from nvidia, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio nvidia 2 variants 744k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What parakeet-tdt is

<span style="color:#76b900;"🦜 parakeet-tdt-0.6b-v3: Multilingual Speech-to-Text Model</span img { display: inline; } [](#model-architecture) <span style="color:#466f00;"Description:</span parakeet-tdt-0.6b-v3 is a 600-million-parameter multilingual automatic speech recognition (ASR) model designed for high-throughput speech-to-text transcription. It extends the parakeet-tdt-0.6b-v2 model by expanding language support from English to 25 European languages. The model automatically detects the language of the audio and transcribes it without requiring additional prompting. It is part of a series of models that leverage the Granary [1, 2] multilingual corpus as their primary training dataset. 🗣️ Try Demo here: https://huggingface.co/spaces/nvidia/parakeet-tdt-0.6b-v3 Supported Languages: Bulgarian (bg), Croatian (hr), Czech (cs), Danish (da), Dutch (nl), English (en), Estonian (et), Finnish (fi), French (fr), German (de), Greek (el), Hungarian (hu), Italian (it), Latvian (lv), Lithuanian (lt), Maltese (mt), Polish (pl), Portuguese (pt), Romanian (ro), Slovak (sk), Slovenian (sl), Spanish (es), Swedish (sv), Russian (ru), Ukrainian (uk) This model is ready for commercial/non-commercial use. <span style="color:#466f00;"Key Features:</span parakeet-tdt-0.6b-v3's key features are built on the foundation of its predecessor, parakeet-tdt-0.6b-v2, and include: Automatic punctuation and capitalization Accurate word-level and segment-level timestamps Long audio transcription, supporting audio up to 24 minutes long with full attention (on A100 80GB) or up to 3 hours with local attention. Released under a permissive CC BY 4.0 license For full details on the model architecture, training methodology, datasets, and evaluation results, check out the Technical Report. <span style="color:#466f00;"License/Terms of Use:</span GOVERNING TERMS: Use of this model is governed by the CC-BY-4.0 license. <span style="color:#466f00;"Discover more from NVIDIA:</span For documentation, deployment guides, enterprise-ready APIs, and the latest open models—including Nemotron and other cutting-edge speech, translation, and generative AI—visit the NVIDIA Developer Portal at developer.nvidia.com. Joi

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makernvidia
TypeAudio & music
Parameters (lead)627M
Variants2
Runs withtransformers
Released2025-08-04
Popularity744k downloads / month
Likes1,538
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
parakeet-tdt-0.6b-v3627MBF16~1.4 GBWeights ↗
parakeet-tdt-0.6b-v2BF16Weights ↗

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
Automatic Speech RecognitionAMI (Meetings test)Test WER11.31
Automatic Speech RecognitionEarnings-22Test WER11.42
Automatic Speech RecognitionGigaSpeechTest WER9.59
Automatic Speech RecognitionLibriSpeech (clean)Test WER1.93
Automatic Speech RecognitionLibriSpeech (clean)Test WER3.59
automatic-speech-recognitionSPGI SpeechTest WER3.97
automatic-speech-recognitiontedlium-v3Test WER2.75
Automatic Speech RecognitionVox PopuliTest WER6.14
automatic-speech-recognitionFLEURSTest WER (Bg)12.64
automatic-speech-recognitionFLEURSTest WER (Cs)11.01
automatic-speech-recognitionFLEURSTest WER (Da)18.41
automatic-speech-recognitionFLEURSTest WER (De)5.04
automatic-speech-recognitionFLEURSTest WER (El)20.7
automatic-speech-recognitionFLEURSTest WER (En)4.85
automatic-speech-recognitionFLEURSTest WER (Es)3.45
automatic-speech-recognitionFLEURSTest WER (Et)17.73
automatic-speech-recognitionFLEURSTest WER (Fi)13.21
automatic-speech-recognitionFLEURSTest WER (Fr)5.15
automatic-speech-recognitionFLEURSTest WER (Hr)12.46
automatic-speech-recognitionFLEURSTest WER (Hu)15.72
automatic-speech-recognitionFLEURSTest WER (It)3
automatic-speech-recognitionFLEURSTest WER (Lt)20.35
automatic-speech-recognitionFLEURSTest WER (Lv)22.84
automatic-speech-recognitionFLEURSTest WER (Mt)20.46

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys nvidia-parakeet-tdt for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (nvidia-parakeet-tdt below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="nvidia-parakeet-tdt" -F file=@audio.mp3

Details

Languages, data & research

Languages

en es fr de bg hr cs da nl et fi el hu it

Trained / evaluated on

nvidia/Granary nemo/asr-set-3.0 nvidia/nemo-asr-set-3.0

Tags

transformers nemo safetensors gguf parakeet_tdt feature-extraction automatic-speech-recognition speech audio Transducer Transformer TDT FastConformer Conformer

Papers

Licence

Open weights

Open weights under cc-by-4.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want parakeet-tdt on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms