Model reference · open weights
Audar-ASR-Flash is an open-weight audio or speech model from audarai. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | audarai |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Parameters (lead) | 782M |
| Runs with | transformers |
| Released | 2026-07-02 |
| Popularity | 1k downloads / month |
| Licence | Commercial licence needed |
About
Audar-ASR-V1-Flash is the edge tier of Audar's Arabic-first speech-recognition family — the same in-house Arabic training program as Audar-ASR-V1-Turbo, delivered in a fast ~0.6B-decoder model for real-time captioning and on-device use. It recasts transcription as audio-conditioned next-token prediction (a language-model decoder, not CTC/transducer), and is built on a permissively-licensed open-weight audio-LLM foundation, then adapted in-house through Audar's Arabic training program — the contribution is the adaptation, not the foundation:
It transcribes MSA and every major Arabic dialect, code-switched Arabic–English, and English, across 30 languages, and runs on CPU / GPU / edge via 🤗 Transformers or GGUF. For maximum accuracy on the hardest dialectal audio, use the larger Turbo tier.
Built on a permissively-licensed open-weight audio-LLM foundation; the adaptation, data, and alignment are Audar's. Full method and results: Audar-ASR-V1 Technical Report. Runs via Transformers, llama.cpp / GGUF, and vLLM.
Flash is evaluated end-to-end on all six leaderboard test sets (full test splits, not sampled), with the leaderboard-equivalent normalizer — the same harness and protocol as every other row (Audar rows below are the leaderboard maintainers' independent reproduction, Aug 2026 normalization). Audar-ASR-V1-Flash scores 32.0 avg WER at just 0.78B parameters — on par with Qwen3-ASR-1.7B (2× its size) and ahead of Voxtral-Small-24B, Whisper-large-v3, and every CTC baseline. Audar's accuracy tier, Turbo, is #1.
Per-dataset WER % across all six sets, plus the two composite averages. Lower is better; Avg WER is the ranking metric. Flash and Turbo (Ours) in bold; bold cell = best in column.
| # | Model | Avg WER | Avg CER | SADA | CV-18 | MASC-clean | MASC-noisy | MGB-2 | Casablanca |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Audar-ASR-V1-Turbo (Ours, 2.35B) | 23.17 | 9.20 | 28.92 | 8.09 | 16.73 | 27.19 | 11.08 | 47.02 |
| 2 | CohereLabs/cohere-transcribe-arabic-07-2026 | 25.87 | 11.80 | 37.47 | 5.82 | 19.60 | 27.07 | 15.54 | 49.71 |
| 3 | omnilingual-asr/omniASR_LLM_7B | 28.32 | 12.52 | 41.61 | 8.75 | 19.69 | 29.29 | 14.13 | 56.46 |
| 4 | omnilingual-asr/omniASR_LLM_3B | 29.96 | 13.77 | 46.18 | 9.15 | 19.90 | 30.03 | 14.22 | 60.27 |
| 5 | omnilingual-asr/omniASR_LLM_1B | 29.96 | 13.40 | 43.84 | 9.55 | 20.03 | 30.26 | 15.34 | 60.68 |
| 6 | CohereLabs/cohere-transcribe-03-2026 | 30.67 | 16.37 | 60.11 | 8.17 | 8.66 | 19.01 | 25.33 | 62.71 |
| 7 | Qwen/Qwen3-Omni-30B-A3B-Instruct | 30.71 | 13.67 | 44.82 | 11.46 | 21.47 | 30.85 | 13.09 | 62.55 |
| 8 | Audar-ASR-V1-Flash (Ours, 0.78B) | 32.04 | 13.43 | 44.36 | 15.38 | 22.56 | 34.19 | 17.05 | 58.71 |
| 9 | nvidia-conformer-ctc-large-arabic (lm) | 32.91 | 13.84 | 44.52 | 8.80 | 23.74 | 34.29 | 17.20 | 68.90 |
| 10 | omnilingual-asr/omniASR_LLM_300M | 32.96 | 14.84 | 51.38 | 12.03 | 20.66 | 32.45 | 16.58 | 64.64 |
| 11 | google/gemma-4-E4B-it | 32.98 | 13.71 | 43.40 | 19.65 | 24.86 | 33.59 | 17.72 | 58.63 |
| 12 | Qwen/Qwen3-ASR-1.7B | 33.36 | 12.33 | 45.53 | 16.90 | 24.37 | 34.29 | 16.57 | 64.47 |
| 13 | mistralai/Voxtral-Small-24B-2507 | 34.47 | 15.29 | 50.82 | 15.25 | 23.96 | 34.43 | 16.03 | 66.30 |
| 14 | nvidia-conformer-ctc-large-arabic (greedy) | 34.74 | 13.37 | 47.26 | 10.60 | 24.12 | 35.64 | 19.69 | 71.13 |
| 15 | google/gemma-4-E2B-it | 35.87 | 15.34 | 46.23 | 23.76 | 27.47 | 36.15 | 20.72 | 60.87 |
| 16 | openai/whisper-large-v3 | 36.86 | 17.21 | 55.96 | 17.83 | 24.66 | 34.63 | 16.26 | 71.81 |
| 17 | omnilingual-asr/omniASR_CTC_3B | 37.78 | 19.79 | 69.85 | 14.19 | 21.48 | 34.60 | 18.96 | 67.58 |
| 18 | omnilingual-asr/omniASR_CTC_7B | 38.12 | 20.91 | 72.69 | 12.47 | 21.08 | 35.04 | 20.43 | 67.02 |
| 19 | facebook/seamless-m4t-v2-large | 38.16 | 17.03 | 62.52 | 21.70 | 25.04 | 33.24 | 20.23 | 66.25 |
| 20 | omnilingual-asr/omniASR_CTC_1B | 39.29 | 20.47 | 71.42 | 17.55 | 22.76 | 35.73 | 19.96 | 68.32 |
| 21 | openai/whisper-large-v3-turbo | 40.05 | 18.87 | 60.36 | 25.73 | 25.51 | 37.16 | 17.75 | 73.79 |
| 22 | openai/whisper-large-v2 | 40.20 | 19.55 | 57.46 | 21.77 | 27.25 | 38.55 | 25.17 | 71.01 |
| 23 | Qwen/Qwen3-ASR-0.6B | 42.19 | 16.23 | 53.75 | 28.28 | 31.34 | 42.63 | 25.45 | 71.68 |
| 24 | openai/whisper-large | 42.57 | 20.49 | 63.24 | 26.04 | 28.89 | 40.79 | 24.28 | 72.18 |
| 25 | mistralai/Voxtral-Mini-3B-2507 | 42.58 | 19.90 | 63.65 | 22.12 | 28.37 | 41.27 | 22.56 | 77.52 |
| 26 | asafaya/hubert-large-arabic-transcribe | 45.50 | 17.35 | 67.82 | 8.01 | 32.94 | 50.16 | 37.51 | 76.53 |
| 27 | openai/whisper-medium | 45.57 | 22.27 | 67.71 | 28.07 | 29.99 | 42.91 | 29.32 | 75.44 |
| 28 | nvidia-Parakeet-ctc-1.1b-concat | 46.54 | 23.88 | 70.70 | 26.34 | 30.49 | 45.95 | 24.94 | 80.80 |
| 29 | omnilingual-asr/omniASR_CTC_300M | 46.65 | 21.86 | 78.11 | 27.90 | 28.40 | 43.26 | 26.85 | 75.35 |
| 30 | nvidia-Parakeet-ctc-1.1b-universal | 51.96 | 25.19 | 73.58 | 40.01 | 36.16 | 50.03 | 30.68 | 81.30 |
| 31 | microsoft/VibeVoice-ASR | 52.99 | 28.95 | 69.83 | 44.25 |
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys audar-asr-flash for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (audar-asr-flash below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \ -H "Authorization: Bearer $AXFORGE_API_KEY" \ -F model="audar-asr-flash" -F file=@audio.mp3
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.