Model reference · open weights
GigaChat3.1-Audio is an open-weight language model from ai-sage. GigaChat3.1-Audio-10B-A1.8B (BF16) weighs 2.4 GB; the smallest configuration that runs it is RTX 3060 12 GB.
GigaChat3.1-Audio is an audio-native large language model developed by ai-sage that integrates a Conformer speech encoder with a Mixture-of-Experts decoder for text generation. It supports Russian and English, operates with a 262,144 token context length, and is released under the MIT license. The model is designed for audio question answering, classification, and temporal grounding tasks such as timestamped event descriptions and summarization.
Summary of the ai-sage/GigaChat3.1-Audio-10B-A1.8B model card, 2026-10-01
What it is
| Released by | ai-sage |
|---|---|
| Type | Language models |
| Task | Text gen |
| Context | 262,144 tokens |
| Runs with | transformers |
| Based on | ai-sage/GigaChat3.1-10B-A1.8B |
| Released | 2026-07-13 |
| Popularity | 280k downloads / month |
| Weights | 2.4 GB (GigaChat3.1-Audio-10B-A1.8B (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 2.4 GB (file size) · KV cache 30 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 938 MB on a small card · context up to 262,144 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB | 33 | 8 | all 256K | 11.6 GB |
| RTX 4060 Ti 16 GB | 49 | 12 | all 256K | 15.4 GB |
| RTX 3090 24 GB | 81 | 20 | all 256K | 23.4 GB |
| RTX 4090 24 GB | 81 | 20 | all 256K | 23.4 GB |
| RTX 5090 32 GB | 112 | 28 | all 256K | 31.0 GB |
| L40S 48 GB | 165 | 41 | all 256K | 44.0 GB |
| A100 80 GB | 305 | 76 | all 256K | 78.2 GB |
| H100 80 GB | 285 | 71 | all 256K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 349 | 87 | all 256K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 404 | 101 | all 256K | 107 GB |
| H200 141 GB | 529 | 132 | all 256K | 138 GB |
| B200 180 GB | 679 | 169 | all 256K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 3.5 GB | 4.3 GB |
| 5 | 4.5 GB | 8.2 GB |
| 8 | 5.3 GB | 11.2 GB |
| 16 | 7.2 GB | 19.0 GB |
| 32 | 11.2 GB | 34.7 GB |
| 64 | 19.0 GB | 66.1 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings. Split over several cards, this model's cache is copied to every card, not divided.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head latent attention (MLA)); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). This model's attention (MLA) copies its cache to every card of a tensor-parallel split; vLLM can also run it data-parallel, each card with its own cache, which holds several times more requests. Assumes vLLM 0.10 or later.
From the model card
GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding speech understanding.
Capabilities: audio question answering and classification, temporal grounding (localization in long audio, timestamped event descriptions, audio summarization with timestamps), tool-use, and text-only tasks.
The temporal grounding skills are trained on TimeGround-1M — a purpose-built dataset of long-form audio paired with time-aligned annotations.
| Task | Set | Metric | GigaChat Audio (10B-A1.8B) | Voxtral (3B) | Phi-4 (4B) | Qwen3-Omni (30B-A3B) |
|---|---|---|---|---|---|---|
| Audio QA | MMAU | acc ↑ | 62.2 | 59.8 | 68.3 | 74.7 |
| Audio QA | MMLU-speech | acc ↑ | 50.3 | 38.8 | 35.1 | 72.2 |
| Audio math | MQA | acc ↑ | 72.5 | 35.3 | 42.0 | 86.7 |
| Audio QA (ru) | RuBQ | acc ↑ | 60.0 | 23.4 | 2.3 | 43.7 |
| Temporal Localization | ≤10m | mIoU ↑ | 40.3 | 3.4 | 0.2 | 12.9 |
| Temporal Localization | 20–60m | mIoU ↑ | 48.3 | 0.1 | 0.2 | 0.1 |
| Emotion | Dusha crowd | acc ↑ | 90.0 | 43.9 | 11.4 | 77.2 |
| Emotion | Dusha podcast | acc ↑ | 92.4 | 79.6 | 7.2 | 80.7 |
| ASR (ru) | Golos crowd | WER ↓ | 14.7 | 25.9 | 180.0 | 13.1 |
| ASR (ru) | Golos farfield | WER ↓ | 9.7 | 30.3 | 188.7 | 18.4 |
| ASR (ru) | FLEURS ru | WER ↓ | 4.4 | 7.8 | 208.5 | 3.3 |
| ASR (en) | FLEURS en | WER ↓ | 6.5 | 4.0 | 4.2 | 5.0 |
| Translation | FLEURS ru→en | BLEU ↑ | 33.4 | 34.0 | 0.1 | 33.8 |
| Translation | FLEURS en→ru | BLEU ↑ | 26.0 | 21.4 | 19.9 | 29.3 |
TL — temporal localization: find when something is discussed (mIoU vs the reference span). TD — timestamped descriptions of audio events (overall grade 0–5). SUM — long-audio summarization with timestamps; single score = mean of factual accuracy, timing structure and audio coverage with timestamps.
| Metric | Bucket | GigaChat Audio (10B-A1.8B) | Voxtral (3B) | Phi-4 (4B) | Qwen3-Omni (30B-A3B) |
|---|---|---|---|---|---|
| TL mIoU ↑ | ≤10m | 40.3 | 3.4 | 0.2 | 12.9 |
| TL mIoU ↑ | 10–20m | 46.8 | 0.0 | 0.2 | 0.0 |
| TL mIoU ↑ | 20–60m | 48.3 | 0.1 | 0.2 | 0.1 |
| TL mIoU ↑ | 60–120m | 48.9 | 0.9 | 3.7 | 4.6 |
| TL mIoU ↑ | AMI meeting | 30.3 | 0.0 | 0.0 | 0.0 |
| TD overall ↑ (0–5) | ≤10m | 3.45 | 2.89 | 2.62 | 3.23 |
| TD overall ↑ (0–5) | 10–20m | 3.21 | 2.33 | 2.07 | 2.49 |
| TD overall ↑ (0–5) | 20–60m | 3.27 | 2.15 | 1.95 | 2.17 |
| TD overall ↑ (0–5) | 60–120m | 3.26 | 1.45 | 1.21 | 1.84 |
| SUM overall ↑ | ≤10m | 71.6 | 64.1 | 26.2 | 65.7 |
| SUM overall ↑ | 10–20m | 71.4 | 58.9 | 23.0 | 54.8 |
| SUM overall ↑ | 20–60m | 67.9 | 50.2 | 21.4 | 43.8 |
| SUM overall ↑ | 60–120m | 55.4 | 40.1 | 9.8 | 17.3 |
Adding the audio modality shifts text quality: some benchmarks regress (MMLU-Pro, IFEval-ru, BBH), others improve (RuBQ, GPQA Diamond).
| Benchmark | Text 10b | Audio 10b | Δ |
|---|---|---|---|
| MMLU_PRO_EN | 62.04 | 52.86 | −9.18 |
| RUBQ (ru) | 67.46 | 68.91 | +1.45 |
| IFEVAL (ru) | 66.22 | 62.35 | −3.87 |
| BBH | 75.72 | 68.46 | −7.26 |
| GPQA Diamond | 39.73 | 40.91 | +1.18 |
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
model_name = "ai-sage/GigaChat3.1-Audio-10B-A1.8B"
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_name, trust_remote_code=True, dtype=torch.bfloat16, device_map="cuda:0",
)
messages = [{"role": "user", "content": [
{"type": "audio", "path": "90min_lecture.wav"},
{"type": "text", "text": "When does the speaker mention Star Wars midi-chlorians?"},
]}]
inputs = processor.prepare_for_inference(messages, device=model.device)
output_ids = model.generate(**inputs, max_new_tokens=256)
answer_ids = output_ids[0, inputs["input_ids"].shape[1]:]
print(processor.decode(answer_ids))
# > ... in the interval 01:27:46 to 01:27:53. In this segment they explain ...
Both encoder and decoder run inside vLLM. compilation_config disables
torch.compile (needed for the Conformer tower); max_num_batched_tokens
sizes the audio encoder cache (~90 min is 33k tokens):
import librosa
import os
from transformers import AutoProcessor
from vllm import LLM, SamplingParams
os.environ.setdefault("VLLM_WORKER_MULTIPROC_METHOD", "spawn")
model_name = "ai-sage/GigaChat3.1-Audio-10B-A1.8B"
def main():
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)
llm = LLM(
model=model_name,
trust_remote_code=True,
max_model_len=65536,
max_num_batched_tokens=65536,
limit_mm_per_prompt={"audio": 1},
compilation_config={"mode": 0, "cudagraph_mode": "FULL"},
)
messages = [{"role": "user", "content": [
{"type": "audio", "path": "90min_lecture.wav"},
{"type": "text", "text": "When does the speaker mention Star Wars midi-chlorians?"},
]}]
text, audio_paths = processor.render_prompt(messages)
wav, _ = librosa.load(audio_paths[0], sr=16000, mono=True)
out = llm.generate(
{"prompt": text, "multi_modal_data": {"audio": [(wav, 16000)]}},
SamplingParams(temperature=0.0, max_tokens=256),
)
print(processor.decode(out[0].outputs[0].token_ids))
# > ... in the interval 01:27:46 to 01:27:53. In this segment they explain ...
if __name__ == "__main__":
main()
Single-request decode speed on one H100 (vLLM 0.18.0, bf16, greedy), by amount of audio held in context:
| In-context audio | Decode throughput (tok/s) |
|---|---|
| ≤ 10 min | 242 |
| 1 |
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.