DiarizationLM-Gemma-4-E4B-v1 Vision-Language Model: 8GB+ VRAM

DiarizationLM-Gemma-4-E4B-v1 Vision-Language Model: 8GB+ VRAM

At a Glance

Item Value
Repository google/DiarizationLM-Gemma-4-E4B-v1
Publisher guide Google: models and licenses
Published 2026-10-04
License apache-2.0
Formats GGUF / safetensors
Paper arXiv:2401.03506
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

Google has released “DiarizationLM-Gemma-4-E4B-v1", a large language model designed for speech processing. This model is based on Google’s “Gemma 4 E4B" and was fine-tuned using the Locality-Preserving Oracle Supervision method. Note that it is explicitly specified as not being an officially supported Google product.

This model takes text output from automatic speech recognition (ASR) and speaker diarization systems, post-processes and corrects turn boundaries and backchannel errors, and outputs optimized text with speaker labels. Unlike traditional models specialized for two-speaker telephone audio, a key feature is that it is optimized across four standard corpora including multi-speaker (up to 9 people) meeting datasets.

Specifications

The published specifications and training conditions are as follows:

  • Base Model: google/gemma-4-E4B (Architecture: Gemma4ForConditionalGeneration, 4B dense parameters / 4.5B active parameters, 8B including embeddings)
  • Task and Input/Output Format: Takes text (ASR hypothesis with speaker tags) as input and outputs corrected text with speaker tags. The prompt format follows the structure <speaker:N> {text} --> {text} [eod]
  • Maximum Sequence Length: Prompt split length of 4,000 characters, maximum sequence length of 2,560 tokens
  • Training Dataset: Total of 71,825 pairs (51,063 items from the Fisher dataset, 20,762 items from Callhome/ICSI/AMI multi-corpus data)
  • Training Optimization Settings: 10,000 steps, global batch size of 8, optimizer is AdamW (beta1 = 0.9, beta2 = 0.99), peak learning rate of 1.5e-4 (with 500 steps of linear warmup and cosine decay applied), conducted on 8 Google Cloud TPU v5p devices
  • License Conditions: Apache 2.0 (an open license allowing commercial use and modification)

Performance and Quality

As evaluation results conducted by the publishers, model evaluation results on four representative diarization benchmarks using USM + turn-to-diarize as a baseline are listed in the model card. Evaluations are calculated using Hungarian matching dynamic programming (diarizationlm.compute_metrics_on_json_dict).

→ Scroll horizontally to see all columns

Corpus Evaluation Set Baseline (USM + Turn-to-Diarize) DiarizationLM-8b-Fisher-v2 (Llama 3 8B) DiarizationLM-Gemma-4-E4B-v1 (4B)
Fisher (2-speaker telephone audio) TEST FULL (172 sessions) 5.32 [4.93, 5.74] 3.28 2.99 [2.65, 3.37]
Callhome (2-5 speaker telephone audio) TEST FULL (20 calls) 7.74 [6.07, 9.65] 6.66 4.92 [3.46, 6.75]
ICSI (3-9 speaker meeting audio) TEST FULL (3 meetings) 14.70 [11.65, 20.29] Not listed 14.10 [10.77, 19.94]
AMI (4-speaker meeting audio) TEST WORD FULL (16 meetings) 15.68 [10.64, 21.11] Not listed 14.89 [9.80, 20.38]

*Values are WDER (Word Diarization Error Rate: lower is better). Numbers in square brackets indicate 95% bootstrap confidence intervals from 10,000 resamplings.

Compared to its predecessor, the Llama 3 8B-based “DiarizationLM-8b-Fisher-v2", this model reduces the error rate to 2.99% on Fisher and 4.92% on Callhome despite having half the parameter count. It also achieves scores significantly lower than the baseline on multi-speaker meeting corpora (ICSI and AMI), which were difficult to evaluate with previous models.

Additionally, benchmark results for the base model “google/gemma-4-E4B" itself have been published as follows:

→ Scroll horizontally to see all columns

Benchmark Gemma 4 31B Gemma 4 26B A4B Gemma 4 12B Unified Gemma 4 E4B Gemma 4 E2B Gemma 3 27B (no think)
MMLU Pro 85.2% 82.6% 77.2% 69.4% 60.0% 67.6%
AIME 2026 no tools 89.2% 88.3% 77.5% 42.5% 37.5% 20.8%
LiveCodeBench v6 80.0% 77.1% 72.0% 52.0% 44.0% 29.1%
Codeforces ELO 2150 1718 1659 940 633 110
GPQA Diamond 84.3% 82.3% 78.8% 58.6% 43.4% 42.4%
Tau2 (average over 3) 76.9% 68.2% 69.0% 42.2% 24.5% 16.2%
BigBench Extra Hard 74.4% 64.8% 53.0% 33.1% 21.9% 19.3%
MMMLU 88.4% 86.3% 83.4% 76.6% 67.4% 70.7%
MMMU Pro 76.9% 73.8% 69.1% 52.6% 44.2% 49.7%
MATH-Vision 85.6% 82.4% 79.7% 59.5% 52.4% 46.0%
CoVoST – – 38.5* 35.54 33.47 –
FLEURS (lower is better) – – 0.069* 0.08 0.09 –

The base model Gemma 4 E4B records 69.4% on MMLU Pro and 58.6% on GPQA Diamond while being a lightweight model, possessing reasoning performance that surpasses the older generation large model Gemma 3 27B (no think). On the other hand, compared to higher-tier models such as the 31B and 26B A4B, it falls behind in advanced reasoning tasks such as difficult mathematics (AIME 2026) and competitive programming (Codeforces).

Strengths and Use Cases

This model excels at post-processing tasks that correct speaker turn boundary judgments and labeling errors with high precision for text output by ASR and existing speaker diarization systems.

Specifically, while it can accurately correct short backchannels of about 1 to 5 words and lexical turn-taking boundaries, it is designed to maintain acoustic speaker anchors during longer utterances (monologues) of 6 or more words, preventing drift phenomena where speaker labels switch midway. This makes it suitable for utilization in dialog and meeting audio processing such as the following:

  • Telephone Conversation Text Optimization: Accurately organizing turn overlaps and backchannels from two-speaker calls (such as Fisher) to casual phone conversations involving roughly 2 to 5 people (such as Callhome).
  • Multi-Participant Meeting Minutes Generation: Suitable for minutes generation and speaker identification in complex environments where multiple people speak actively, such as four-person face-to-face meetings (AMI) or academic research meetings with up to 9 participants (ICSI).
  • Accuracy Improvement of Existing ASR Pipelines: Improving the Word Diarization Error Rate (WDER) simply by adding format conversion of output text and lightweight inference, without needing to retrain existing speech recognition models or diarization processes.

Note that while the base model “google/gemma-4-E4B" is a multimodal foundational model supporting native image and audio processing, this derivative model “DiarizationLM-Gemma-4-E4B-v1" is fine-tuned specifically for text processing tasks of taking text-input ASR hypotheses and generating optimized text with speaker labels.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 8.0B parameters

Your VRAM Quantization File size Est. memory needed
8GB (RTX 4060 / 3060 Ti, etc.) Q4_K_M 4.9GB 5.9GB

Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-05): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

Runs in Ollama, LM Studio and llama.cpp as-is.

It is distributed in GGUF, so no conversion is needed.

License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.

Compression: the Q4_K_M build measures 5.26 bits per weight — about 33% the size of the original 16-bit weights, calculated by this site from the actual file sizes.

Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Distributed Files

Weight files published in google/DiarizationLM-Gemma-4-E4B-v1, listed by this site from the Hugging Face API. Sizes are the actual file sizes.

File Size
DiarizationLM-Gemma-4-E4B-v1-q4_0.gguf 5.15GB
DiarizationLM-Gemma-4-E4B-v1-q4_k_m.gguf 5.30GB
model.safetensors 15.99GB

How to Get It

This model is available in the Hugging Face repository “google/DiarizationLM-Gemma-4-E4B-v1". Since it is not a gated model, it can be directly downloaded and used without waiting for additional usage requests or prior approvals.

Weight files are provided in the following formats:

  • safetensors: A 16-bit (bfloat16) format that can be loaded directly with Hugging Face’s transformers.
  • GGUF: The recommended 4-bit K-quant Medium (Q4_K_M) and legacy 4-bit (Q4_0) formats are available. The recommended Q4_K_M version adopts 256-element superblocks while keeping sensitive layers like attn_v, ffn_down, and embedding matrices in 6-bit (Q6_K).

Execution Steps in Python (transformers + diarizationlm)

When using in a Python environment, perform GPU inference after installing the related libraries. By combining it with the dedicated library diarizationlm, post-processing to reflect LLM output results back into the original hypothesis text can be smoothly executed.

First, install the necessary packages.

pip install transformers diarizationlm

Next, run model loading and inference with the following code.

from diarizationlm import utils
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "google/DiarizationLM-Gemma-4-E4B-v1"

HYPOTHESIS = (
    "<speaker:1> Hello, how are you doing <speaker:2> today? I am doing well."
    " What about <speaker:1> you? I'm doing well, too. Thank you."
)

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, device_map="cuda")
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID, torch_dtype=torch.bfloat16, device_map="cuda"
)

inputs = tokenizer([HYPOTHESIS + " --> "], return_tensors="pt").to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=int(inputs.input_ids.shape[1] * 1.2),
    do_sample=False,
    use_cache=True,
)

completion = tokenizer.batch_decode(
    outputs[:, inputs.input_ids.shape[1]:], skip_special_tokens=True
)[0]
completion = utils.truncate_suffix_and_tailing_text(completion, " [eod]")

transferred_completion = utils.transfer_llm_completion(completion, HYPOTHESIS)

print("Hypothesis:", HYPOTHESIS)
print("Transferred completion:", transferred_completion)

Execution Steps in llama.cpp

It is also possible to run the GGUF version directly using inference engines like llama.cpp, Ollama, or llama-cpp-python. When using llama-cli, execute the command as follows:

llama-cli \
  -m DiarizationLM-Gemma-4-E4B-v1-q4_k_m.gguf \
  -p "<speaker:1> Hello, how are you doing <speaker:2> today? I am doing well. What about <speaker:1> you? I'm doing well, too. Thank you. --> " \
  --temp 0.0 \
  -n 128

Other Models for the Same Task

Recent vision-language models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site’s estimate.

See all vision-language models →

What to Read Next

Sources