Model reference · open weights

DiarizationLM-Gemma-4

NEW · this week LLMs google Vision + text 1 build Open weights 0 dl/mo

DiarizationLM-Gemma-4 is an open-weight language model from Google. DiarizationLM-Gemma-4-E4B-v1 (BF16) weighs 5.3 GB; the smallest configuration that runs it is RTX 3060 12 GB.

  • DiarizationLM-Gemma-4 is a large language model developed by Google to post-process and correct automatic speech recognition and speaker diarization outputs.
  • It is built on the Gemma 4 E4B foundation model with 4 billion dense parameters and supports a context length of 131,072 tokens.
  • The model is trained on four canonical speaker diarization benchmark corpora and is released under the Apache 2.0 license.

Summary of the google/DiarizationLM-Gemma-4-E4B-v1 model card, 2026-10-05

What it is

Released byGoogle
Released2026-10-04
Parameters8.0B
VRAM5.3 GB for the weights

What it runs on

Memory and cards for DiarizationLM-Gemma-4-E4B-v1 (BF16)

5.3 GBweights, file size
14 MBcache per 1K tokens
330 MBwindow cache per request
901 MBruntime overhead, at least
131,072 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB126all 128K11.6 GB
RTX 4060 Ti 16 GB2011all 128K15.4 GB
RTX 3090 24 GB3821all 128K23.4 GB
RTX 4090 24 GB3821all 128K23.4 GB
RTX 5090 32 GB5531all 128K31.0 GB
L40S 48 GB8447all 128K44.0 GB
A100 80 GB16089all 128K78.2 GB
H100 80 GB10242all 128K78.1 GB
RTX PRO 6000 Blackwell 96 GB12452all 128K93.8 GB
DGX Spark (GB10) 128 GB unified14360all 128K107 GB
H200 141 GB18778all 128K138 GB
B200 180 GB24160all 128K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
16.6 GB7.0 GB
58.4 GB10.2 GB
89.8 GB12.6 GB
1613.4 GB19.0 GB
3220.5 GB31.8 GB
6434.9 GB57.4 GB

One card, with vLLM's small-card settings.

From the model card

What Google says about DiarizationLM-Gemma-4

Read the model card

This is not an officially supported Google product.

Overview

DiarizationLM is a Large Language Model framework designed to post-process, correct, and optimize automatic speech recognition (ASR) and speaker diarization outputs.

google/DiarizationLM-Gemma-4-E4B-v1 is built on Google's Gemma 4 E4B (4B dense parameters) foundation model and fine-tuned with Locality-Preserving Oracle Supervision across all 4 canonical speaker diarization benchmark corpora:

  1. Fisher English (2-speaker conversational telephone speech)
  2. Callhome American English (2–5 speaker informal telephone conversations)
  3. ICSI Meeting Corpus (3–9 speaker academic research meetings)
  4. AMI Meeting Corpus (4-speaker tabletop meetings)
  • Foundation model: google/gemma-4-E4B
  • Open-source library & scripts: https://github.com/google/speaker-id/tree/master/DiarizationLM

Unlike earlier models trained exclusively on 2-speaker telephone data (such as google/DiarizationLM-8b-Fisher-v2) or naive multi-domain SFT (which suffers from long-monologue speaker identity drift in 4+ speaker meetings), google/DiarizationLM-Gemma-4-E4B-v1 is trained on locality-preserving oracle targets that teach the model to correct lexical turn boundaries and backchannels (1 ≤ L ≤ 5 words) while preserving acoustic speaker anchors across long monologues (L ≥ 6 words). As a result, it achieves statistically significant (p < 0.0001) WDER and cpWER improvements simultaneously across all four benchmarks under standard out-of-the-box Transcript-Preserving Speaker Transfer (transfer_llm_completion), despite having half the parameter count (4B vs. 8B).

Training Configuration

  • Base Architecture: Gemma 4 E4B (42 layers, hidden size 2560, hybrid 5:1 sliding-window and global attention, 262,144 vocabulary size)
  • LoRA Adapter: Rank r = 256 applied to all attention (q_proj, k_proj, v_proj, o_proj), MLP (gate_proj, up_proj, down_proj), and per-layer input (per_layer_input_gate, per_layer_projection, per_layer_model_projection) linear projections, merged into 16-bit (bfloat16) base weights and serialized to 4-bit GGUF (Q4_K_M and Q4_0)
  • Training Objective: Completion-only cross-entropy loss ( --> [eod])
  • Training Data: 51,063 Fisher + 20,762 Multi-Corpus (Callhome, ICSI, AMI) locality-preserving prompt-completion pairs
  • Optimization: 10,000 steps, global batch size 8, AdamW (beta1 = 0.9, beta2 = 0.99), peak learning rate 1.5e-4 with 500-step linear warmup and cosine decay on 8 Google Cloud TPU v5p chips
  • Prompt Segmentation Length: 4,000 characters (maximal sequence length 2,560 tokens)

Model Files Included

  • model.safetensors: Merged 16-bit (bfloat16) Hugging Face transformers weights (~16.0 GB)
  • DiarizationLM-Gemma-4-E4B-v1-q4_k_m.gguf: (Recommended GGUF) Serialized 4-bit K-quant Medium (Q4_K_M) GGUF model (~5.30 GB) for llama.cpp / Ollama / llama-cpp-python (uses 256-element super-blocks in Q4_K with sensitive attn_v / ffn_down and embedding matrices retained in 6-bit Q6_K)
  • DiarizationLM-Gemma-4-E4B-v1-q4_0.gguf: Serialized legacy 4-bit (Q4_0) GGUF model (~5.15 GB) for llama.cpp / Ollama / llama-cpp-python
  • config.json, generation_config.json, tokenizer.json, tokenizer_config.json, processor_config.json, chat_template.jinja, special_tokens_map.json: Tokenizer and model configuration files

Benchmark Performance (with 95% Bootstrap Confidence Intervals)

All metrics below are micro-averaged across the full evaluation sets using the USM + turn-to-diarize baseline and scored via Hungarian-matching dynamic programming (diarizationlm.compute_metrics_on_json_dict). Ranges in brackets indicate 95% non-parametric bootstrap confidence intervals (B = 10,000 conversation-level resamples):

BenchmarkTest SplitWER (%)SystemWDER (%) [95% CI]cpWER (%) [95% CI]SpkCntMAE [95% CI]Paired ΔWDER vs. Baseline (p-value)
FisherTEST FULL (172 sessions)15.37Baseline (USM + Turn-to-Diarize)DiarizationLM-8b-Fisher-v2 (Llama 3 8B)DiarizationLM-Gemma-4-E4B-v1 (4B)5.32 [4.93, 5.74]3.282.99 [2.65, 3.37]20.88 [19.97, 21.88]18.3717.62 [16.76, 18.56]0.215 [0.151, 0.291]---0.093 [0.047, 0.145]-------2.33% [-2.51, -2.16] (p < 0.0001)
CallhomeTEST FULL (20 calls)15.22Baseline (USM + Turn-to-Diarize)DiarizationLM-8b-Fisher-v2 (Llama 3 8B)DiarizationLM-Gemma-4-E4B-v1 (4B)7.74 [6.07, 9.65]6.664.92 [3.46, 6.75]24.31 [21.27, 27.43]23.5720.69 [17.98, 23.66]0.050 [0.000, 0.150]---0.000 [0.000, 0.000]-------2.82% [-3.48, -2.14] (p < 0.0001)
ICSITEST FULL (3 meetings)29.42Baseline (USM + Turn-to-Diarize)DiarizationLM-Gemma-4-E4B-v1 (4B)14.70 [11.65, 20.29]14.10 [10.77, 19.94]43.90 [39.10, 51.67]43.32 [38.38, 51.20]0.333 [0.000, 1.000]0.333 [0.000, 1.000]----0.60% [-0.88, -0.35] (p < 0.0001)
AMITEST WORD FULL (16 meetings)24.33Baseline (USM + Turn-to-Diarize)DiarizationLM-Gemma-4-E4B-v1 (4B)15.68 [10.64, 21.11]14.89 [9.80, 20.38]40.32 [32.43, 48.00]39.57 [31.57, 47.29]0.500 [0.188, 0.812]0.438 [0.188, 0.750]----0.79% [-1.00, -0.60] (p < 0.0001)

Usage

1. Python (transformers + diarizationlm)

First, install the required packages:

pip install transformers diarizationlm

Run inference on a GPU with bfloat16:

from diarizationlm import utils
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "google/DiarizationLM-Gemma-4-E4B-v1"

HYPOTHESIS = (
    " Hello, how are you doing  today? I am doing well."
    " What about  you? I'm do

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms