Model reference · open weights
Audio8-ASR-Infinite is an open-weight audio or speech model from Edge0. Audio8-ASR-Infinite (BF16) weighs 8.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.
Audio8-ASR-Infinite is a 4.1B parameter automatic speech recognition model developed by Edge0 for streaming Chinese and English transcription. It features a 32,768 token context length and supports unlimited-length audio processing with constant memory and latency. The model is released under the Apache 2.0 license.
Summary of the Edge0/Audio8-ASR-Infinite model card, 2026-10-01
What it is
| Released by | Edge0 |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Parameters (lead) | 4.1B |
| Context | 32,768 tokens |
| Runs with | transformers |
| Released | 2026-09-21 |
| Popularity | 2k downloads / month |
| Weights | 8.2 GB (Audio8-ASR-Infinite (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 8.2 GB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
Audio8 ASR Infinite is a native streaming speech recognition model built to be as responsive as possible. It offers a selectable audio clock (80/120/160 ms) and a transcription delay (240–560 ms). With our adapted vLLM build it transcribes unlimited-length audio 24/7 without drifting.
The checkpoint has a native context of 30 seconds. But with Rolling KV Cache, it can transcribe 24/7 nonstop.
src="https://huggingface.co/Edge0/Audio8-ASR-Infinite/resolve/main/Audio8-Asr-Infinite-Demo.mp4">
The following combinations of frame length and delay are post-trained. Other combinations can be used but performance may not be optimum.
| audio clock | frame_len | streaming_n_left_pad_tokens | selectable target_delay_ms |
|---|---|---|---|
| 80 ms | 4 | 18 | 240 / 320 / 480 / 560 |
| 120 ms | 6 | 12 | 240 / 480 |
| 160 ms | 8 | 9 | 320 / 480 |
target_delay_ms must be an integer multiple of the selected clock, so longer
delays stay available at every clock even when they are not listed above.
Inherits the Voxtral realtime audio architecture and DSM-style streaming.
| Component | Initial weights | Trained |
|---|---|---|
| Causal Audio Tower | Voxtral Realtime 4B | ✅ |
| Audio Projector | random initialisation | ✅ |
| Frame Length Embedding | random initialisation | ✅ |
| Decoder | Qwen2.5-3B-Instruct | ✅ |
| LM Head | Qwen2.5-3B-Instruct | ✅ |
Checkpoint specification:
| audio tower | 32 layers, hidden 1280, 128 mel bins, sliding window 750 |
| text decoder | 36 layers, hidden 2048, 16 query heads / 2 KV heads |
| projector | max frame len 8 → projection size 10240, gelu |
| frame-length conditioning | enabled (use_frame_len_embedding: true) |
| semantic VAD heads | semantic_vad_heads.safetensors, 8 classes, horizons 0.5 / 1.0 / 2.0 / 3.0 s |
| vocab size | 151936 |
| dtype | bfloat16 |
| weights | 8.17 GB model.safetensors (+ semantic_vad_heads.safetensors) |
This is the preview release: it delivers the transcription base. Realtime semantic perception is being built on the same frame grid and the same acoustic forward pass.
| Stage | Status | Scope |
|---|---|---|
| Preview — ASR base | ✅ done | Streaming Chinese/English transcription: selectable 80/120/160 ms clock, configurable target_delay_ms, unlimited-length rolling KV window |
| Formal release | 🏃in progress | Frame-level semantic perception on the same grid, beyond transcription |
| test set | metric | Audio8 ASR Infinite | Voxtral-Mini-4B-Realtime-2602 | nemotron-3.5-asr-streaming-0.6b |
|---|---|---|---|---|
| aishell1/test | CER | 1.750 | 16.795 | 12.927@560ms |
| aishell4/test | CER | 2.893 | 16.456 | 14.677@560ms |
| librispeech test.clean | WER | 3.042 | 2.210 | 3.353@560ms |
| librispeech test.other | WER | 6.808 | 5.552 | 7.140@560ms |
| average | 3.623 | 10.253 (2 sets) | 9.524 |
Greedy decode with EOS suppressed, at the 80 ms audio clock with
target_delay_ms = 480 (6 delay tokens). Error rates in percent. No repetition
loops and no dropped trailing words.
Programmatic simulated-streaming decode with the embedded remote code:
import numpy as np
import torch
from transformers import AutoFeatureExtractor, AutoTokenizer
from audio8_asr_infinite.modeling.modeling_audio8_asr_infinite import (
Audio8ASRInfiniteForConditionalGeneration,
resolve_qwen_language_token_id,
resolve_qwen_streaming_special_token_ids,
)
from audio8_asr_infinite.streaming_inference import simulated_streaming_greedy_decode_batch
checkpoint = "Edge0/Audio8-ASR-Infinite"
tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True)
feature_extractor = AutoFeatureExtractor.from_pretrained(checkpoint, trust_remote_code=True)
model = Audio8ASRInfiniteForConditionalGeneration.from_pretrained(
checkpoint, trust_remote_code=True, torch_dtype=torch.bfloat16
).eval().cuda()
class AudioConfig: # duck-typed: raw_audio_samples_per_token / streaming_n_left_pad_tokens / sampling_rate
raw_audio_samples_per_token = 1280 # 80 ms @ 16 kHz
streaming_n_left_pad_tokens = 18
sampling_rate = 16000
waveform = np.load("sample.npy", allow_pickle=False).astype(np.float32) # [-1, 1], 16 kHz mono
results = simulated_streaming_greedy_decode_batch(
model=model,
tokenizer=tokenizer,
feature_extractor=feature_extractor,
waveforms=[waveform],
language_token_ids=[resolve_qwen_language_token_id(tokenizer, "zh")],
special_ids=resolve_qwen_streaming_special_token_ids(tokenizer),
audio_config=AudioConfig(),
num_delay_tokens=[480 // 80],
right_pad_text_tokens=10,
dtype=torch.bfloat16,
device=next(model.parameters()).device,
max_new_tokens=512,
)
print(results[0]["final_text"])
Only a full merged weight directory is supported (this repository as-is); adapter-style or partially converted weights are not.
Docker compos
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.