Model reference · open weights

GigaChat3.1-Audio

LLMs ai-sage Text gen 1 build Open weights 280k dl/mo

GigaChat3.1-Audio is an open-weight language model from ai-sage. GigaChat3.1-Audio-10B-A1.8B (BF16) weighs 2.4 GB; the smallest configuration that runs it is RTX 3060 12 GB.

GigaChat3.1-Audio is an audio-native large language model developed by ai-sage that integrates a Conformer speech encoder with a Mixture-of-Experts decoder for text generation. It supports Russian and English, operates with a 262,144 token context length, and is released under the MIT license. The model is designed for audio question answering, classification, and temporal grounding tasks such as timestamped event descriptions and summarization.

Summary of the ai-sage/GigaChat3.1-Audio-10B-A1.8B model card, 2026-10-01

What it is

Released byai-sage
TypeLanguage models
TaskText gen
Context262,144 tokens
Runs withtransformers
Based onai-sage/GigaChat3.1-10B-A1.8B
Released2026-07-13
Popularity280k downloads / month
Weights2.4 GB (GigaChat3.1-Audio-10B-A1.8B (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for GigaChat3.1-Audio-10B-A1.8B (BF16)

Weights 2.4 GB (file size) · KV cache 30 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 938 MB on a small card · context up to 262,144 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB338all 256K11.6 GB
RTX 4060 Ti 16 GB4912all 256K15.4 GB
RTX 3090 24 GB8120all 256K23.4 GB
RTX 4090 24 GB8120all 256K23.4 GB
RTX 5090 32 GB11228all 256K31.0 GB
L40S 48 GB16541all 256K44.0 GB
A100 80 GB30576all 256K78.2 GB
H100 80 GB28571all 256K78.1 GB
RTX PRO 6000 Blackwell 96 GB34987all 256K93.8 GB
DGX Spark (GB10) 128 GB unified404101all 256K107 GB
H200 141 GB529132all 256K138 GB
B200 180 GB679169all 256K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
13.5 GB4.3 GB
54.5 GB8.2 GB
85.3 GB11.2 GB
167.2 GB19.0 GB
3211.2 GB34.7 GB
6419.0 GB66.1 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings. Split over several cards, this model's cache is copied to every card, not divided.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head latent attention (MLA)); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). This model's attention (MLA) copies its cache to every card of a tensor-parallel split; vLLM can also run it data-parallel, each card with its own cache, which holds several times more requests. Assumes vLLM 0.10 or later.

From the model card

What ai-sage says about GigaChat3.1-Audio

Read the model card

GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding speech understanding.

Capabilities: audio question answering and classification, temporal grounding (localization in long audio, timestamped event descriptions, audio summarization with timestamps), tool-use, and text-only tasks.

The temporal grounding skills are trained on TimeGround-1M — a purpose-built dataset of long-form audio paired with time-aligned annotations.

Evaluation

1. Core audio tasks vs open models

TaskSetMetricGigaChat Audio (10B-A1.8B)Voxtral (3B)Phi-4 (4B)Qwen3-Omni (30B-A3B)
Audio QAMMAUacc ↑62.259.868.374.7
Audio QAMMLU-speechacc ↑50.338.835.172.2
Audio mathMQAacc ↑72.535.342.086.7
Audio QA (ru)RuBQacc ↑60.023.42.343.7
Temporal Localization≤10mmIoU ↑40.33.40.212.9
Temporal Localization20–60mmIoU ↑48.30.10.20.1
EmotionDusha crowdacc ↑90.043.911.477.2
EmotionDusha podcastacc ↑92.479.67.280.7
ASR (ru)Golos crowdWER ↓14.725.9180.013.1
ASR (ru)Golos farfieldWER ↓9.730.3188.718.4
ASR (ru)FLEURS ruWER ↓4.47.8208.53.3
ASR (en)FLEURS enWER ↓6.54.04.25.0
TranslationFLEURS ru→enBLEU ↑33.434.00.133.8
TranslationFLEURS en→ruBLEU ↑26.021.419.929.3

2. Timing tasks (detailed)

TL — temporal localization: find when something is discussed (mIoU vs the reference span). TD — timestamped descriptions of audio events (overall grade 0–5). SUM — long-audio summarization with timestamps; single score = mean of factual accuracy, timing structure and audio coverage with timestamps.

MetricBucketGigaChat Audio (10B-A1.8B)Voxtral (3B)Phi-4 (4B)Qwen3-Omni (30B-A3B)
TL mIoU ↑≤10m40.33.40.212.9
TL mIoU ↑10–20m46.80.00.20.0
TL mIoU ↑20–60m48.30.10.20.1
TL mIoU ↑60–120m48.90.93.74.6
TL mIoU ↑AMI meeting30.30.00.00.0
TD overall ↑ (0–5)≤10m3.452.892.623.23
TD overall ↑ (0–5)10–20m3.212.332.072.49
TD overall ↑ (0–5)20–60m3.272.151.952.17
TD overall ↑ (0–5)60–120m3.261.451.211.84
SUM overall ↑≤10m71.664.126.265.7
SUM overall ↑10–20m71.458.923.054.8
SUM overall ↑20–60m67.950.221.443.8
SUM overall ↑60–120m55.440.19.817.3

3. Text quality vs the text base model

Adding the audio modality shifts text quality: some benchmarks regress (MMLU-Pro, IFEval-ru, BBH), others improve (RuBQ, GPQA Diamond).

BenchmarkText 10bAudio 10bΔ
MMLU_PRO_EN62.0452.86−9.18
RUBQ (ru)67.4668.91+1.45
IFEVAL (ru)66.2262.35−3.87
BBH75.7268.46−7.26
GPQA Diamond39.7340.91+1.18

Quickstart (Transformers)

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

model_name = "ai-sage/GigaChat3.1-Audio-10B-A1.8B"
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_name, trust_remote_code=True, dtype=torch.bfloat16, device_map="cuda:0",
)

messages = [{"role": "user", "content": [
    {"type": "audio", "path": "90min_lecture.wav"},
    {"type": "text", "text": "When does the speaker mention Star Wars midi-chlorians?"},
]}]
inputs = processor.prepare_for_inference(messages, device=model.device)
output_ids = model.generate(**inputs, max_new_tokens=256)
answer_ids = output_ids[0, inputs["input_ids"].shape[1]:]
print(processor.decode(answer_ids))
# > ... in the interval 01:27:46 to 01:27:53. In this segment they explain ...

vLLM (single GPU, native multimodal)

Both encoder and decoder run inside vLLM. compilation_config disables torch.compile (needed for the Conformer tower); max_num_batched_tokens sizes the audio encoder cache (~90 min is 33k tokens):

import librosa
import os
from transformers import AutoProcessor
from vllm import LLM, SamplingParams

os.environ.setdefault("VLLM_WORKER_MULTIPROC_METHOD", "spawn")

model_name = "ai-sage/GigaChat3.1-Audio-10B-A1.8B"

def main():
    processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)
    llm = LLM(
        model=model_name,
        trust_remote_code=True,
        max_model_len=65536,
        max_num_batched_tokens=65536,
        limit_mm_per_prompt={"audio": 1},
        compilation_config={"mode": 0, "cudagraph_mode": "FULL"},
    )

    messages = [{"role": "user", "content": [
        {"type": "audio", "path": "90min_lecture.wav"},
        {"type": "text", "text": "When does the speaker mention Star Wars midi-chlorians?"},
    ]}]
    text, audio_paths = processor.render_prompt(messages)
    wav, _ = librosa.load(audio_paths[0], sr=16000, mono=True)
    out = llm.generate(
        {"prompt": text, "multi_modal_data": {"audio": [(wav, 16000)]}},
        SamplingParams(temperature=0.0, max_tokens=256),
    )
    print(processor.decode(out[0].outputs[0].token_ids))
    # > ... in the interval 01:27:46 to 01:27:53. In this segment they explain ...

if __name__ == "__main__":
    main()

Throughput

Single-request decode speed on one H100 (vLLM 0.18.0, bf16, greedy), by amount of audio held in context:

In-context audioDecode throughput (tok/s)
≤ 10 min242
1

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms