Model reference · open weights

FireRedASR2-LLM-vllm

Available as managed deployment Audio allendou · community Speech→text 1 variants 530 dl/mo

FireRedASR2-LLM-vllm is an open-weight audio or speech model from allendou. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byallendou
TypeAudio & music
TaskSpeech→text
Context32k tokens
Released2026-02-26
Popularity530 downloads / month
LicenceOpen weights

About

What FireRedASR2-LLM-vllm is

FireRedASR2S A SOTA Industrial-Grade All-in-One ASR System

[Code] [Paper] [Model] [Blog] [Demo]

FireRedASR2S is a state-of-the-art (SOTA), industrial-grade, all-in-one ASR system with ASR, VAD, LID, and Punc modules. All modules achieve SOTA performance:

Read the full model card
  • FireRedASR2: Automatic Speech Recognition (ASR) supporting Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and singing lyrics recognition. 2.89% average CER on Mandarin (4 test sets), 11.55% on Chinese dialects (19 test sets), outperforming Doubao-ASR, Qwen3-ASR-1.7B, Fun-ASR, and Fun-ASR-Nano-2512. FireRedASR2-AED also supports word-level timestamps and confidence scores.
  • FireRedVAD: Voice Activity Detection (VAD) supporting speech/singing/music in 100+ languages. 97.57% F1, outperforming Silero-VAD, TEN-VAD, and FunASR-VAD. Supports non-streaming/streaming VAD and Audio Event Detection.
  • FireRedLID: Spoken Language Identification (LID) supporting 100+ languages and 20+ Chinese dialects/accents. 97.18% accuracy, outperforming Whisper and SpeechBrain-LID.
  • FireRedPunc: Punctuation Prediction (Punc) for Chinese and English. 78.90% average F1, outperforming FunASR-Punc (62.77%).

2S: 2nd-generation FireRedASR, now expanded to an all-in-one ASR System

🔥 News

  • [2026.02.25] 🔥 We release FireRedASR2-LLM model weights. 🤗 🤖
  • [2026.02.13] 🚀 Support TensorRT-LLM inference acceleration for FireRedASR2-AED (contributed by NVIDIA). Benchmark on AISHELL-1 test set shows 12.7x speedup over PyTorch baseline (single H20).
  • [2026.02.12] 🔥 We release FireRedASR2S (FireRedASR2-AED, FireRedVAD, FireRedLID, and FireRedPunc) with model weights and inference code. Download links below. Technical report and finetuning code coming soon.

Available Models and Languages

ModelSupported Languages & DialectsDownload
FireRedASR2-LLMChinese (Mandarin and 20+ dialects/accents*), English, Code-Switching🤗 | 🤖
FireRedASR2-AEDChinese (Mandarin and 20+ dialects/accents*), English, Code-Switching🤗 | 🤖
FireRedVAD100+ languages, 20+ Chinese dialects/accents*🤗 | 🤖
FireRedLID100+ languages, 20+ Chinese dialects/accents*🤗 | 🤖
FireRedPuncChinese, English🤗 | 🤖

Method

FireRedASR2

FireRedASR2 builds upon FireRedASR with improved accuracy, designed to meet diverse requirements in superior performance and optimal efficiency across various applications. It comprises two variants:

  • FireRedASR2-LLM: Designed to achieve state-of-the-art performance and to enable seamless end-to-end speech interaction. It adopts an Encoder-Adapter-LLM framework leveraging large language model (LLM) capabilities.
  • FireRedASR2-AED: Designed to balance high performance and computational efficiency and to serve as an effective speech representation module in LLM-based speech models. It utilizes an Attention-based Encoder-Decoder (AED) architecture.

Other Modules

  • FireRedVAD: DFSMN-based non-streaming/streaming Voice Activity Detection and Audio Event Detection.
  • FireRedLID: FireRedASR2-based Spoken Language Identification. See FireRedLID README for language details.
  • FireRedPunc: BERT-based Punctuation Prediction.

Evaluation

FireRedASR2

Metrics: Character Error Rate (CER%) for Chinese and Word Error Rate (WER%) for English. Lower is better.

We evaluate FireRedASR2 on 24 public test sets covering Mandarin, 20+ Chinese dialects/accents, and singing.

  • Mandarin (4 test sets): 2.89% (LLM) / 3.05% (AED) average CER, outperforming Doubao-ASR (3.69%), Qwen3-ASR-1.7B (3.76%), Fun-ASR (4.16%) and Fun-ASR-Nano-2512 (4.55%).
  • Dialects (19 test sets): 11.55% (LLM) / 11.67% (AED) average CER, outperforming Doubao-ASR (15.39%), Qwen3-ASR-1.7B (11.85%), Fun-ASR (12.76%) and Fun-ASR-Nano-2512 (15.07%).

Note: ws=WenetSpeech, md=MagicData, conv=Conversational, daily=Daily-use.

IDTestset\ModelFireRedASR2-LLMFireRedASR2-AEDDoubao-ASRQwen3-ASRFun-ASRFun-ASR-Nano
Average CER(All, 1-24)9.679.8012.9810.1210.9212.81
Average CER(Mandarin, 1-4)2.893.053.693.764.164.55
Average CER(Dialects, 5-23)11.5511.6715.3911.8512.7615.07
1aishell10.640.571.521.481.641.96
2aishell22.152.512.772.712.383.02
3ws-net4.444.575.734.976.856.93
4ws-meeting4.324.534.745.885.786.29
5kespeech3.083.605.385.105.367.66
6ws-yue-short5.145.1510.515.827.348.82
7ws-yue-long8.718.5411.398.8510.1411.36
8ws-chuan-easy10.9010.6011.3311.9912.4614.05
9ws-chuan-hard20.7121.3520.7721.6322.4925.32
10md-heavy7.427.437.698.029.139.97
11md-yue-conv12.2311.6626.259.7633.7115.68
12

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys fireredasr2-llm-vllm for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (fireredasr2-llm-vllm below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="fireredasr2-llm-vllm" -F file=@audio.mp3

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms