Model reference · open weights

canary

Available as managed deployment Audio nvidia Speech→text 1 variants 12k dl/mo

canary is an open-weight audio or speech model from nvidia. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Makernvidia
TypeAudio & music
TaskSpeech→text
Parameters (lead)979M
Runs withnemo
Released2025-08-04
Popularity12k downloads / month
LicenceOpen weights

About

What canary is

🐤 Canary 1B v2: Multitask Speech Transcription and Translation Model

Canary-1b-v2 is a powerful 1-billion parameter model built for high-quality speech transcription and translation across 25 European languages.

It excels at both automatic speech recognition (ASR) and speech translation (AST), supporting:

  • Speech Transcription (ASR) for 25 languages
  • Speech Translation (AST) from English → 24 languages
  • Speech Translation (AST) from 24 languages → English

Supported Languages: Bulgarian (bg), Croatian (hr), Czech (cs), Danish (da), Dutch (nl), English (en), Estonian (et), Finnish (fi), French (fr), German (de), Greek (el), Hungarian (hu), Italian (it), Latvian (lv), Lithuanian (lt), Maltese (mt), Polish (pl), Portuguese (pt), Romanian (ro), Slovak (sk), Slovenian (sl), Spanish (es), Swedish (sv), Russian (ru), Ukrainian (uk)

🗣️ Experience Canary-1b-v2 in action at Hugging Face Demo

Canary-1b-v2 model is ready for commercial/non-commercial use.

License/Terms of Use:

GOVERNING TERMS: Use of this model is governed by the CC-BY-4.0 license.

Discover more from NVIDIA:

For documentation, deployment guides, enterprise-ready APIs, and the latest open models—including Nemotron and other cutting-edge speech, translation, and generative AI—visit the NVIDIA Developer Portal at developer.nvidia.com. Join the community to access tools, support, and resources to accelerate your development with NVIDIA’s NeMo, Riva, NIM, and foundation models.

Explore more from NVIDIA:

What is Nemotron? NVIDIA Developer Nemotron NVIDIA Riva Speech NeMo Documentation

Key Features

Canary-1b-v2 is a scaled and enhanced version of the Canary model family, offering:

  • Support for 25 European languages, expanding from the 4 languages in canary-1b/canary-1b-flash to 21 additional languages
  • State-of-the-art performance among models of similar size
  • Comparable quality to models 3× larger, while being up to 10× faster
  • Automatic punctuation and capitalization
  • Accurate word-level and segment-level timestamps
  • Segment-level timestamps also available for translated outputs
  • Released under a permissive CC BY 4.0 license

Canary-1b-v2 model is the first model from NeMo team that leveraged full Nvidia's Granary dataset [1] [2], showcasing its multitask and multilingual capabilities.

For full details on the model architecture, training methodology, datasets, and evaluation results, check out the Canary-1b-v2 Technical Report.

For a deeper glimpse into the Canary family of models, explore this comprehensive NeMo tutorial on multitask speech models.

Automatic Speech Recognition (ASR)

Figure 1: ASR WER comparison across different models. This does not include Punctuation and Capitalisation errors.


Speech Translation (AST)

X → English

Figure 2: AST X → En COMET scores comparison across different models

English → X

Figure 3: AST En → X COMET scores comparison across different models


Evaluation Notes

Note 1: The above evaluations are conducted in two settings: (1) All supported languages (24 languages, excluding Latvian since seamless-m4t-v2-large and seamless-m4t-medium do not support it), and (2) Common languages (6 languages supported by all compared models: en, fr, de, it, pt, es).

Note 2: Performance differences may be partly attributed to Portuguese variant differences - our training data uses European Portuguese while most benchmarks use Brazilian Portuguese.


Deployment Geography

Global

Use case

This model serves developers, researchers, academics, and industries building applications that require speech-to-text capabilities, including but not limited to: conversational AI, voice assistants, transcription services, subtitle generation, and voice analytics platforms.

Release Date

Huggingface 08/14/2025

Model Architecture

Canary-1b-v2 is an encoder-decoder architecture featuring a FastConformer Encoder [3] and a Transformer Decoder [4]. The model extracts audio features through the encoder and uses task-specific tokens—such as and—to guide the Transformer Decoder in generating text output.

It uses a unified SentencePiece Tokenizer [5] with a vocabulary of 16,384 tokens, optimized across all 25 supported languages. The architecture includes 32 encoder layers and 8 decoder layers, totaling 978 million parameters.

For implementation details, see the NeMo repository.

Input

  • Input Type(s): 16kHz Audio
  • Input Format(s): .wav and .flac audio formats
  • Input Parameters: 1D (audio signal)
  • Other Properties Related to Input: Monochannel audio

Output

  • Output Type(s): Text
  • Output Format: String
  • Output Parameters: 1D (text)
  • Other Properties Related to Output: Punctuation and Capitalization included.

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faste

From the published model card. Full card on the HuggingFace links in the sidebar.

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
automatic-speech-recognitionFLEURSTest WER (Bg)9.250
automatic-speech-recognitionFLEURSTest WER (Cs)7.860
automatic-speech-recognitionFLEURSTest WER (Da)11.250
automatic-speech-recognitionFLEURSTest WER (De)4.400
automatic-speech-recognitionFLEURSTest WER (El)9.210
automatic-speech-recognitionFLEURSTest WER (En)4.500
automatic-speech-recognitionFLEURSTest WER (Es)2.900
automatic-speech-recognitionFLEURSTest WER (Et)12.550
automatic-speech-recognitionFLEURSTest WER (Fi)8.590
automatic-speech-recognitionFLEURSTest WER (Fr)5.020
automatic-speech-recognitionFLEURSTest WER (Hr)8.290
automatic-speech-recognitionFLEURSTest WER (Hu)12.900
automatic-speech-recognitionFLEURSTest WER (It)3.070
automatic-speech-recognitionFLEURSTest WER (Lt)12.360
automatic-speech-recognitionFLEURSTest WER (Lv)9.660
automatic-speech-recognitionFLEURSTest WER (Mt)18.310
automatic-speech-recognitionFLEURSTest WER (Nl)6.120
automatic-speech-recognitionFLEURSTest WER (Pl)6.640
automatic-speech-recognitionFLEURSTest WER (Pt)4.390
automatic-speech-recognitionFLEURSTest WER (Ro)6.610
automatic-speech-recognitionFLEURSTest WER (Ru)6.900
automatic-speech-recognitionFLEURSTest WER (Sk)5.740
automatic-speech-recognitionFLEURSTest WER (Sl)13.320
automatic-speech-recognitionFLEURSTest WER (Sv)9.570

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys canary for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (canary below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="canary" -F file=@audio.mp3

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms