Model reference · open weights

Audio8-ASR-Infinite

Audio Edge0 Speech→text 1 build Open weights 2k dl/mo

Audio8-ASR-Infinite is an open-weight audio or speech model from Edge0. Audio8-ASR-Infinite (BF16) weighs 8.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.

Audio8-ASR-Infinite is a 4.1B parameter automatic speech recognition model developed by Edge0 for streaming Chinese and English transcription. It features a 32,768 token context length and supports unlimited-length audio processing with constant memory and latency. The model is released under the Apache 2.0 license.

Summary of the Edge0/Audio8-ASR-Infinite model card, 2026-10-01

What it is

Released byEdge0
TypeAudio & music
TaskSpeech→text
Parameters (lead)4.1B
Context32,768 tokens
Runs withtransformers
Released2026-09-21
Popularity2k downloads / month
Weights8.2 GB (Audio8-ASR-Infinite (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for Audio8-ASR-Infinite (BF16)

Weights 8.2 GB (file size) · overhead about 1.6 GB.

CardOne streamCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What Edge0 says about Audio8-ASR-Infinite

Read the model card

Audio8 ASR Infinite is a native streaming speech recognition model built to be as responsive as possible. It offers a selectable audio clock (80/120/160 ms) and a transcription delay (240–560 ms). With our adapted vLLM build it transcribes unlimited-length audio 24/7 without drifting.

Highlights

  • Super responsive — the native streaming architecture decodes 12.5 times per second.
  • Unlimited-length transcription — a rolling KV Cache keeps both memory and latency constant, even in 24/7 operation.
  • Selectable streaming clock — one text token per clock step (12.5 / 8.3 / 6.25 decisions per second), balancing perception granularity and resource cost.
  • Configurable transcription delay — set how much delay to trade for accuracy.
  • Semantic VAD — distinguishes thinking pauses, stuttering and real end of turn, where traditional acoustic VAD usually fails.
  • Bilingual — Chinese and English.

See Audio8-ASR-Infinite in action

The checkpoint has a native context of 30 seconds. But with Rolling KV Cache, it can transcribe 24/7 nonstop.

src="https://huggingface.co/Edge0/Audio8-ASR-Infinite/resolve/main/Audio8-Asr-Infinite-Demo.mp4">

Optimized operation points

The following combinations of frame length and delay are post-trained. Other combinations can be used but performance may not be optimum.

audio clockframe_lenstreaming_n_left_pad_tokensselectable target_delay_ms
80 ms418240 / 320 / 480 / 560
120 ms612240 / 480
160 ms89320 / 480

target_delay_ms must be an integer multiple of the selected clock, so longer delays stay available at every clock even when they are not listed above.

Architecture

Inherits the Voxtral realtime audio architecture and DSM-style streaming.

ComponentInitial weightsTrained
Causal Audio TowerVoxtral Realtime 4B✅
Audio Projectorrandom initialisation✅
Frame Length Embeddingrandom initialisation✅
DecoderQwen2.5-3B-Instruct✅
LM HeadQwen2.5-3B-Instruct✅

Checkpoint specification:

audio tower32 layers, hidden 1280, 128 mel bins, sliding window 750
text decoder36 layers, hidden 2048, 16 query heads / 2 KV heads
projectormax frame len 8 → projection size 10240, gelu
frame-length conditioningenabled (use_frame_len_embedding: true)
semantic VAD headssemantic_vad_heads.safetensors, 8 classes, horizons 0.5 / 1.0 / 2.0 / 3.0 s
vocab size151936
dtypebfloat16
weights8.17 GB model.safetensors (+ semantic_vad_heads.safetensors)

Roadmap

This is the preview release: it delivers the transcription base. Realtime semantic perception is being built on the same frame grid and the same acoustic forward pass.

StageStatusScope
Preview — ASR base✅ doneStreaming Chinese/English transcription: selectable 80/120/160 ms clock, configurable target_delay_ms, unlimited-length rolling KV window
Formal release🏃in progressFrame-level semantic perception on the same grid, beyond transcription

Evaluation

480 ms Delay, 80ms frame length

test setmetricAudio8 ASR InfiniteVoxtral-Mini-4B-Realtime-2602nemotron-3.5-asr-streaming-0.6b
aishell1/testCER1.75016.79512.927@560ms
aishell4/testCER2.89316.45614.677@560ms
librispeech test.cleanWER3.0422.2103.353@560ms
librispeech test.otherWER6.8085.5527.140@560ms
average3.62310.253 (2 sets)9.524

Greedy decode with EOS suppressed, at the 80 ms audio clock with target_delay_ms = 480 (6 delay tokens). Error rates in percent. No repetition loops and no dropped trailing words.

Usage

Programmatic simulated-streaming decode with the embedded remote code:

import numpy as np
import torch
from transformers import AutoFeatureExtractor, AutoTokenizer

from audio8_asr_infinite.modeling.modeling_audio8_asr_infinite import (
    Audio8ASRInfiniteForConditionalGeneration,
    resolve_qwen_language_token_id,
    resolve_qwen_streaming_special_token_ids,
)
from audio8_asr_infinite.streaming_inference import simulated_streaming_greedy_decode_batch

checkpoint = "Edge0/Audio8-ASR-Infinite"
tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True)
feature_extractor = AutoFeatureExtractor.from_pretrained(checkpoint, trust_remote_code=True)
model = Audio8ASRInfiniteForConditionalGeneration.from_pretrained(
    checkpoint, trust_remote_code=True, torch_dtype=torch.bfloat16
).eval().cuda()

class AudioConfig:  # duck-typed: raw_audio_samples_per_token / streaming_n_left_pad_tokens / sampling_rate
    raw_audio_samples_per_token = 1280   # 80 ms @ 16 kHz
    streaming_n_left_pad_tokens = 18
    sampling_rate = 16000

waveform = np.load("sample.npy", allow_pickle=False).astype(np.float32)  # [-1, 1], 16 kHz mono
results = simulated_streaming_greedy_decode_batch(
    model=model,
    tokenizer=tokenizer,
    feature_extractor=feature_extractor,
    waveforms=[waveform],
    language_token_ids=[resolve_qwen_language_token_id(tokenizer, "zh")],
    special_ids=resolve_qwen_streaming_special_token_ids(tokenizer),
    audio_config=AudioConfig(),
    num_delay_tokens=[480 // 80],
    right_pad_text_tokens=10,
    dtype=torch.bfloat16,
    device=next(model.parameters()).device,
    max_new_tokens=512,
)
print(results[0]["final_text"])

Only a full merged weight directory is supported (this repository as-is); adapter-style or partially converted weights are not.

24/7 inference with vLLM

Docker compos

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms