Model reference · open weights

Audar-ASR-Flash

Available as managed deployment Licence fee Audio audarai Speech→text 1 variants 1k dl/mo

Audar-ASR-Flash is an open-weight audio or speech model from audarai. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byaudarai
TypeAudio & music
TaskSpeech→text
Parameters (lead)782M
Runs withtransformers
Released2026-07-02
Popularity1k downloads / month
LicenceCommercial licence needed

About

What Audar-ASR-Flash is

Audar's Arabic-first ASR — the real-time, edge tier.

From Arabic to the world.

Read the full model card

🧭 What it is

Audar-ASR-V1-Flash is the edge tier of Audar's Arabic-first speech-recognition family — the same in-house Arabic training program as Audar-ASR-V1-Turbo, delivered in a fast ~0.6B-decoder model for real-time captioning and on-device use. It recasts transcription as audio-conditioned next-token prediction (a language-model decoder, not CTC/transducer), and is built on a permissively-licensed open-weight audio-LLM foundation, then adapted in-house through Audar's Arabic training program — the contribution is the adaptation, not the foundation:

  • 🧱 Large-scale bilingual pretraining — 300,000+ hours of labeled audio, primarily Arabic and English (MSA + Gulf, Egyptian, Levantine, Maghrebi; code-switching; diverse channels).
  • 🎯 Dialect-targeted fine-tuning with hardness and multi-task sampling.
  • 🧠 KTO preference alignment (Kahneman-Tversky Optimization) from trained native-Arabic annotators.

It transcribes MSA and every major Arabic dialect, code-switched Arabic–English, and English, across 30 languages, and runs on CPU / GPU / edge via 🤗 Transformers or GGUF. For maximum accuracy on the hardest dialectal audio, use the larger Turbo tier.

Built on a permissively-licensed open-weight audio-LLM foundation; the adaptation, data, and alignment are Audar's. Full method and results: Audar-ASR-V1 Technical Report. Runs via Transformers, llama.cpp / GGUF, and vLLM.

Model summary

📊 Benchmarks

Open Universal Arabic ASR Leaderboard — full standings

Flash is evaluated end-to-end on all six leaderboard test sets (full test splits, not sampled), with the leaderboard-equivalent normalizer — the same harness and protocol as every other row (Audar rows below are the leaderboard maintainers' independent reproduction, Aug 2026 normalization). Audar-ASR-V1-Flash scores 32.0 avg WER at just 0.78B parameters — on par with Qwen3-ASR-1.7B (2× its size) and ahead of Voxtral-Small-24B, Whisper-large-v3, and every CTC baseline. Audar's accuracy tier, Turbo, is #1.

Per-dataset WER % across all six sets, plus the two composite averages. Lower is better; Avg WER is the ranking metric. Flash and Turbo (Ours) in bold; bold cell = best in column.

#ModelAvg WERAvg CERSADACV-18MASC-cleanMASC-noisyMGB-2Casablanca
1Audar-ASR-V1-Turbo (Ours, 2.35B)23.179.2028.928.0916.7327.1911.0847.02
2CohereLabs/cohere-transcribe-arabic-07-202625.8711.8037.475.8219.6027.0715.5449.71
3omnilingual-asr/omniASR_LLM_7B28.3212.5241.618.7519.6929.2914.1356.46
4omnilingual-asr/omniASR_LLM_3B29.9613.7746.189.1519.9030.0314.2260.27
5omnilingual-asr/omniASR_LLM_1B29.9613.4043.849.5520.0330.2615.3460.68
6CohereLabs/cohere-transcribe-03-202630.6716.3760.118.178.6619.0125.3362.71
7Qwen/Qwen3-Omni-30B-A3B-Instruct30.7113.6744.8211.4621.4730.8513.0962.55
8Audar-ASR-V1-Flash (Ours, 0.78B)32.0413.4344.3615.3822.5634.1917.0558.71
9nvidia-conformer-ctc-large-arabic (lm)32.9113.8444.528.8023.7434.2917.2068.90
10omnilingual-asr/omniASR_LLM_300M32.9614.8451.3812.0320.6632.4516.5864.64
11google/gemma-4-E4B-it32.9813.7143.4019.6524.8633.5917.7258.63
12Qwen/Qwen3-ASR-1.7B33.3612.3345.5316.9024.3734.2916.5764.47
13mistralai/Voxtral-Small-24B-250734.4715.2950.8215.2523.9634.4316.0366.30
14nvidia-conformer-ctc-large-arabic (greedy)34.7413.3747.2610.6024.1235.6419.6971.13
15google/gemma-4-E2B-it35.8715.3446.2323.7627.4736.1520.7260.87
16openai/whisper-large-v336.8617.2155.9617.8324.6634.6316.2671.81
17omnilingual-asr/omniASR_CTC_3B37.7819.7969.8514.1921.4834.6018.9667.58
18omnilingual-asr/omniASR_CTC_7B38.1220.9172.6912.4721.0835.0420.4367.02
19facebook/seamless-m4t-v2-large38.1617.0362.5221.7025.0433.2420.2366.25
20omnilingual-asr/omniASR_CTC_1B39.2920.4771.4217.5522.7635.7319.9668.32
21openai/whisper-large-v3-turbo40.0518.8760.3625.7325.5137.1617.7573.79
22openai/whisper-large-v240.2019.5557.4621.7727.2538.5525.1771.01
23Qwen/Qwen3-ASR-0.6B42.1916.2353.7528.2831.3442.6325.4571.68
24openai/whisper-large42.5720.4963.2426.0428.8940.7924.2872.18
25mistralai/Voxtral-Mini-3B-250742.5819.9063.6522.1228.3741.2722.5677.52
26asafaya/hubert-large-arabic-transcribe45.5017.3567.828.0132.9450.1637.5176.53
27openai/whisper-medium45.5722.2767.7128.0729.9942.9129.3275.44
28nvidia-Parakeet-ctc-1.1b-concat46.5423.8870.7026.3430.4945.9524.9480.80
29omnilingual-asr/omniASR_CTC_300M46.6521.8678.1127.9028.4043.2626.8575.35
30nvidia-Parakeet-ctc-1.1b-universal51.9625.1973.5840.0136.1650.0330.6881.30
31microsoft/VibeVoice-ASR52.9928.9569.8344.25

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys audar-asr-flash for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (audar-asr-flash below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="audar-asr-flash" -F file=@audio.mp3

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms