Model reference · open weights
qwen3-asr-onnx is an open-weight audio or speech model from rhasspy. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Maker | rhasspy |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Runs with | onnxruntime |
| Based on | andrewleech/qwen3-asr-0.6b-onnx |
| Released | 2026-08-12 |
| Popularity | 100 downloads / month |
| Licence | Open weights |
About
andrewleech/qwen3-asr-0.6b-onnx
with the audio encoder requantized to int4. The decoders, embedding table,
config, and tokenizer are that repo's published int4 files, unchanged.
Ultimately derived from Qwen/Qwen3-ASR-0.6B.
Apache-2.0 throughout.
The upstream repo's encoder.int4.onnx is misleadingly named: it holds FP32
weights (746 MB). Quantizing it for real shrinks the package, and on ARM it cuts
latency too.
| upstream | this | |
|---|---|---|
| encoder | 746 MB (FP32) | 121 MB (int4) |
| total on disk | 2.0 GB | 1.4 GB |
| Raspberry Pi 5, 3.5 s utterance, 4 threads | 1.577 s | 1.36–1.49 s |
| Ryzen 9 5950X, 3.5 s utterance, 16 threads | 0.408 s | 0.411 s |
| Pi 5 peak RSS, short utterance | — | 1.7 GB |
The encoder is compute-bound, so the speedup is ARM-only; on x86 this is purely a memory saving. Transcripts matched the FP32 encoder on 9 of 10 Home Assistant test clips — the tenth was an entity name both get wrong without a context prompt and both get right with one.
| File | Description |
|---|---|
encoder.int4.onnx (+ .data) | Audio encoder, int4 MatMulNBits |
decoder_init.int4.onnx | Decoder prefill; takes input_ids, emits logits + KV cache |
decoder_step.int4.onnx | Autoregressive step; takes input_embeds + KV cache |
decoder_weights.int4.data | Shared external weights for both decoders |
embed_tokens.bin | Token embeddings [151936, 1024], float16 |
config.json, tokenizer.json | Architecture config, special tokens, mel params, tokenizer |
Keep embed_tokens.bin in fp16 and cast one row per generated token; casting the
whole table at load costs ~300 MB of RSS for nothing.
encoder.int4.onnx → audio featuresdecoder_init.int4.onnx with input_ids,
position_ids, audio_features, audio_offsetdecoder_step.int4.onnx until orThe prompt is the Qwen chat template:
Note the token ids for the words system and user are 8948 and 872.
The upstream reference src/prompt.py hardcodes 9125 and 882, which decode to
" Current" and " time".
Free-form text in the system turn biases decoding toward specific spellings — useful for smart-home entity names:
system: Vocabulary: Ecobee, office lamp.
"What's the temperature of the incubator?" → "What's the temperature of the Ecobee?"
Verified clean with a 77-token entity list: no dropped outputs, no prompt echo.
(The int8 export at OpenVoiceOS/qwen3-asr-0.6b-onnx degenerates under the same
load — empty output, or echoing the vocabulary list back as the transcript. int4
MatMulNBits' per-group scales handle the decoder's outlier weights that per-tensor
int8 does not.)
import onnx
from onnxruntime.quantization.matmul_nbits_quantizer import (
MatMulNBitsQuantizer, RTNWeightOnlyQuantConfig)
from onnxruntime.quantization.quant_utils import QuantFormat
q = MatMulNBitsQuantizer(
model=onnx.load("encoder.int4.onnx"), # the FP32-weighted file from upstream
block_size=64, is_symmetric=False, accuracy_level=4,
algo_config=RTNWeightOnlyQuantConfig(quant_format=QuantFormat.QOperator),
)
q.process()
q.model.save_model_to_file("encoder.int4.onnx", use_external_data_format=True)
block_size=64 / accuracy_level=4 match the recipe the decoders were built
with. Don't change them casually — upstream measured block_size=32 at the same
accuracy level producing 99.98% WER, and requantizing the decoder with these
settings produced empty output on both x86 and ARM.
From the published model card. Full card on the HuggingFace links in the sidebar.
How it works
Using it via the API
Once AxForge deploys qwen3-asr-onnx for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (qwen3-asr-onnx below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \ -H "Authorization: Bearer $AXFORGE_API_KEY" \ -F model="qwen3-asr-onnx" -F file=@audio.mp3
Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.