Model reference · open weights
DiarizationLM-Gemma-4 is an open-weight language model from Google. DiarizationLM-Gemma-4-E4B-v1 (BF16) weighs 5.3 GB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the google/DiarizationLM-Gemma-4-E4B-v1 model card, 2026-10-05
What it is
| Released by | |
|---|---|
| Released | 2026-10-04 |
| Parameters | 8.0B |
| VRAM | 5.3 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB | 12 | 6 | all 128K | 11.6 GB |
| RTX 4060 Ti 16 GB | 20 | 11 | all 128K | 15.4 GB |
| RTX 3090 24 GB | 38 | 21 | all 128K | 23.4 GB |
| RTX 4090 24 GB | 38 | 21 | all 128K | 23.4 GB |
| RTX 5090 32 GB | 55 | 31 | all 128K | 31.0 GB |
| L40S 48 GB | 84 | 47 | all 128K | 44.0 GB |
| A100 80 GB | 160 | 89 | all 128K | 78.2 GB |
| H100 80 GB | 102 | 42 | all 128K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 124 | 52 | all 128K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 143 | 60 | all 128K | 107 GB |
| H200 141 GB | 187 | 78 | all 128K | 138 GB |
| B200 180 GB | 241 | 60 | all 128K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 6.6 GB | 7.0 GB |
| 5 | 8.4 GB | 10.2 GB |
| 8 | 9.8 GB | 12.6 GB |
| 16 | 13.4 GB | 19.0 GB |
| 32 | 20.5 GB | 31.8 GB |
| 64 | 34.9 GB | 57.4 GB |
One card, with vLLM's small-card settings.
From the model card
This is not an officially supported Google product.
DiarizationLM is a Large Language Model framework designed to post-process, correct, and optimize automatic speech recognition (ASR) and speaker diarization outputs.
google/DiarizationLM-Gemma-4-E4B-v1 is built on Google's Gemma 4 E4B (4B dense parameters) foundation model and fine-tuned with Locality-Preserving Oracle Supervision across all 4 canonical speaker diarization benchmark corpora:
Unlike earlier models trained exclusively on 2-speaker telephone data (such as google/DiarizationLM-8b-Fisher-v2) or naive multi-domain SFT (which suffers from long-monologue speaker identity drift in 4+ speaker meetings), google/DiarizationLM-Gemma-4-E4B-v1 is trained on locality-preserving oracle targets that teach the model to correct lexical turn boundaries and backchannels (1 ≤ L ≤ 5 words) while preserving acoustic speaker anchors across long monologues (L ≥ 6 words). As a result, it achieves statistically significant (p < 0.0001) WDER and cpWER improvements simultaneously across all four benchmarks under standard out-of-the-box Transcript-Preserving Speaker Transfer (transfer_llm_completion), despite having half the parameter count (4B vs. 8B).
q_proj, k_proj, v_proj, o_proj), MLP (gate_proj, up_proj, down_proj), and per-layer input (per_layer_input_gate, per_layer_projection, per_layer_model_projection) linear projections, merged into 16-bit (bfloat16) base weights and serialized to 4-bit GGUF (Q4_K_M and Q4_0) --> [eod])model.safetensors: Merged 16-bit (bfloat16) Hugging Face transformers weights (~16.0 GB)DiarizationLM-Gemma-4-E4B-v1-q4_k_m.gguf: (Recommended GGUF) Serialized 4-bit K-quant Medium (Q4_K_M) GGUF model (~5.30 GB) for llama.cpp / Ollama / llama-cpp-python (uses 256-element super-blocks in Q4_K with sensitive attn_v / ffn_down and embedding matrices retained in 6-bit Q6_K)DiarizationLM-Gemma-4-E4B-v1-q4_0.gguf: Serialized legacy 4-bit (Q4_0) GGUF model (~5.15 GB) for llama.cpp / Ollama / llama-cpp-pythonconfig.json, generation_config.json, tokenizer.json, tokenizer_config.json, processor_config.json, chat_template.jinja, special_tokens_map.json: Tokenizer and model configuration filesAll metrics below are micro-averaged across the full evaluation sets using the USM + turn-to-diarize baseline and scored via Hungarian-matching dynamic programming (diarizationlm.compute_metrics_on_json_dict). Ranges in brackets indicate 95% non-parametric bootstrap confidence intervals (B = 10,000 conversation-level resamples):
| Benchmark | Test Split | WER (%) | System | WDER (%) [95% CI] | cpWER (%) [95% CI] | SpkCntMAE [95% CI] | Paired ΔWDER vs. Baseline (p-value) |
|---|---|---|---|---|---|---|---|
| Fisher | TEST FULL (172 sessions) | 15.37 | Baseline (USM + Turn-to-Diarize)DiarizationLM-8b-Fisher-v2 (Llama 3 8B)DiarizationLM-Gemma-4-E4B-v1 (4B) | 5.32 [4.93, 5.74]3.282.99 [2.65, 3.37] | 20.88 [19.97, 21.88]18.3717.62 [16.76, 18.56] | 0.215 [0.151, 0.291]---0.093 [0.047, 0.145] | -------2.33% [-2.51, -2.16] (p < 0.0001) |
| Callhome | TEST FULL (20 calls) | 15.22 | Baseline (USM + Turn-to-Diarize)DiarizationLM-8b-Fisher-v2 (Llama 3 8B)DiarizationLM-Gemma-4-E4B-v1 (4B) | 7.74 [6.07, 9.65]6.664.92 [3.46, 6.75] | 24.31 [21.27, 27.43]23.5720.69 [17.98, 23.66] | 0.050 [0.000, 0.150]---0.000 [0.000, 0.000] | -------2.82% [-3.48, -2.14] (p < 0.0001) |
| ICSI | TEST FULL (3 meetings) | 29.42 | Baseline (USM + Turn-to-Diarize)DiarizationLM-Gemma-4-E4B-v1 (4B) | 14.70 [11.65, 20.29]14.10 [10.77, 19.94] | 43.90 [39.10, 51.67]43.32 [38.38, 51.20] | 0.333 [0.000, 1.000]0.333 [0.000, 1.000] | ----0.60% [-0.88, -0.35] (p < 0.0001) |
| AMI | TEST WORD FULL (16 meetings) | 24.33 | Baseline (USM + Turn-to-Diarize)DiarizationLM-Gemma-4-E4B-v1 (4B) | 15.68 [10.64, 21.11]14.89 [9.80, 20.38] | 40.32 [32.43, 48.00]39.57 [31.57, 47.29] | 0.500 [0.188, 0.812]0.438 [0.188, 0.750] | ----0.79% [-1.00, -0.60] (p < 0.0001) |
transformers + diarizationlm)First, install the required packages:
pip install transformers diarizationlm
Run inference on a GPU with bfloat16:
from diarizationlm import utils
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_ID = "google/DiarizationLM-Gemma-4-E4B-v1"
HYPOTHESIS = (
" Hello, how are you doing today? I am doing well."
" What about you? I'm doQuoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.