Model reference · open weights

GigaAM

GigaAM is an open-weight audio or speech model from ai-sage, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio ai-sage 1 variants 232k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What GigaAM is

GigaAM-v3 GigaAM-v3 is a Conformer-based foundation model with 220–240M parameters, pretrained on diverse Russian speech data using the HuBERT-CTC objective. It is the third generation of the GigaAM family and provides state-of-the-art performance on Russian ASR across a wide range of domains. GigaAM-v3 includes the following model variants: - ssl — self-supervised HuBERT–CTC encoder pre-trained on 700,000 hours of Russian speech - ctc — ASR model fine-tuned with a CTC decoder - rnnt — ASR model fine-tuned with an RNN-T decoder - e2ectc — end-to-end CTC model with punctuation and text normalization - e2ernnt — end-to-end RNN-T model with punctuation and text normalization GigaAM-v3 training incorporates new internal datasets: callcenter conversations, speech with background music, natural speech, and speech with atypical characteristics. the models perform on average 30% better on these new domains, while maintaining the same quality as previous GigaAM generations on public benchmarks. The table below reports the Word Error Rate (%) for GigaAM-v3 and other existing models over diverse domains. The end-to-end ASR models (e2ectc and e2ernnt) produce punctuated, normalized text directly. In end-to-end ASR comparisons of e2ectc and e2ernnt against Whisper-large-v3, using Gemini 2.5 Pro as an LLM-as-a-judge, GigaAM-v3 models win by an average margin of 70:30. For detailed results, see metrics. Usage Recommended versions: - torch==2.8.0, torchaudio==2.8.0 - transformers==4.57.1 - pyannote-audio==4.0.0, torchcodec==0.7.0 - (any) hydra-core, omegaconf, sentencepiece Full usage guide can be found in the example. License: MIT Paper: GigaAM: Efficient Self-Supervised Learner for Speech Recognition (InterSpeech 2025)

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makerai-sage
TypeAudio & music
Variants1
Released2025-11-19
Popularity232k downloads / month
Likes145
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
GigaAM-v3BF16Weights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys gigaam for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (gigaam below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="gigaam" -F file=@audio.mp3

Details

Languages, data & research

Languages

ru en

Tags

pytorch gigaam automatic-speech-recognition custom_code ru en

Papers

Licence

Open weights

Open weights under mit — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want GigaAM on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms