{"id":9817,"date":"2026-10-05T01:09:04","date_gmt":"2026-10-04T16:09:04","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/05\/diarizationlm-gemma-4-e4b-v1\/"},"modified":"2026-10-05T01:09:04","modified_gmt":"2026-10-04T16:09:04","slug":"diarizationlm-gemma-4-e4b-v1","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/05\/diarizationlm-gemma-4-e4b-v1\/","title":{"rendered":"DiarizationLM-Gemma-4-E4B-v1 Vision-Language Model: 8GB+ VRAM"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/google\/DiarizationLM-Gemma-4-E4B-v1\">google\/DiarizationLM-Gemma-4-E4B-v1<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-google-en\/\">Google: models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-04<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Paper<\/td>\n<td><a href=\"https:\/\/arxiv.org\/abs\/2401.03506\">arXiv:2401.03506<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>Google has released &#8220;DiarizationLM-Gemma-4-E4B-v1&#8221;, a large language model designed for speech processing. This model is based on Google&#8217;s &#8220;Gemma 4 E4B&#8221; and was fine-tuned using the Locality-Preserving Oracle Supervision method. Note that it is explicitly specified as not being an officially supported Google product.<\/p>\n<p>This model takes text output from automatic speech recognition (ASR) and speaker diarization systems, post-processes and corrects turn boundaries and backchannel errors, and outputs optimized text with speaker labels. Unlike traditional models specialized for two-speaker telephone audio, a key feature is that it is optimized across four standard corpora including multi-speaker (up to 9 people) meeting datasets.<\/p>\n<h2>Specifications<\/h2>\n<p>The published specifications and training conditions are as follows:<\/p>\n<ul>\n<li><strong>Base Model<\/strong>: google\/gemma-4-E4B (Architecture: Gemma4ForConditionalGeneration, 4B dense parameters \/ 4.5B active parameters, 8B including embeddings)<\/li>\n<li><strong>Task and Input\/Output Format<\/strong>: Takes text (ASR hypothesis with speaker tags) as input and outputs corrected text with speaker tags. The prompt format follows the structure <code>&lt;speaker:N&gt; {text} --&gt; {text} [eod]<\/code><\/li>\n<li><strong>Maximum Sequence Length<\/strong>: Prompt split length of 4,000 characters, maximum sequence length of 2,560 tokens<\/li>\n<li><strong>Training Dataset<\/strong>: Total of 71,825 pairs (51,063 items from the Fisher dataset, 20,762 items from Callhome\/ICSI\/AMI multi-corpus data)<\/li>\n<li><strong>Training Optimization Settings<\/strong>: 10,000 steps, global batch size of 8, optimizer is AdamW (beta1 = 0.9, beta2 = 0.99), peak learning rate of 1.5e-4 (with 500 steps of linear warmup and cosine decay applied), conducted on 8 Google Cloud TPU v5p devices<\/li>\n<li><strong>License Conditions<\/strong>: Apache 2.0 (an open license allowing commercial use and modification)<\/li>\n<\/ul>\n<h2>Performance and Quality<\/h2>\n<p>As evaluation results conducted by the publishers, model evaluation results on four representative diarization benchmarks using USM + turn-to-diarize as a baseline are listed in the model card. Evaluations are calculated using Hungarian matching dynamic programming (<code>diarizationlm.compute_metrics_on_json_dict<\/code>).<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th style=\"text-align: left;\">Corpus<\/th>\n<th style=\"text-align: left;\">Evaluation Set<\/th>\n<th style=\"text-align: left;\">Baseline (USM + Turn-to-Diarize)<\/th>\n<th style=\"text-align: left;\">DiarizationLM-8b-Fisher-v2 (Llama 3 8B)<\/th>\n<th style=\"text-align: left;\">DiarizationLM-Gemma-4-E4B-v1 (4B)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"text-align: left;\"><strong>Fisher<\/strong> (2-speaker telephone audio)<\/td>\n<td style=\"text-align: left;\">TEST FULL (172 sessions)<\/td>\n<td style=\"text-align: left;\">5.32 [4.93, 5.74]<\/td>\n<td style=\"text-align: left;\">3.28<\/td>\n<td style=\"text-align: left;\"><strong>2.99 [2.65, 3.37]<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\"><strong>Callhome<\/strong> (2-5 speaker telephone audio)<\/td>\n<td style=\"text-align: left;\">TEST FULL (20 calls)<\/td>\n<td style=\"text-align: left;\">7.74 [6.07, 9.65]<\/td>\n<td style=\"text-align: left;\">6.66<\/td>\n<td style=\"text-align: left;\"><strong>4.92 [3.46, 6.75]<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\"><strong>ICSI<\/strong> (3-9 speaker meeting audio)<\/td>\n<td style=\"text-align: left;\">TEST FULL (3 meetings)<\/td>\n<td style=\"text-align: left;\">14.70 [11.65, 20.29]<\/td>\n<td style=\"text-align: left;\">Not listed<\/td>\n<td style=\"text-align: left;\"><strong>14.10 [10.77, 19.94]<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\"><strong>AMI<\/strong> (4-speaker meeting audio)<\/td>\n<td style=\"text-align: left;\">TEST WORD FULL (16 meetings)<\/td>\n<td style=\"text-align: left;\">15.68 [10.64, 21.11]<\/td>\n<td style=\"text-align: left;\">Not listed<\/td>\n<td style=\"text-align: left;\"><strong>14.89 [9.80, 20.38]<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>*Values are WDER (Word Diarization Error Rate: lower is better). Numbers in square brackets indicate 95% bootstrap confidence intervals from 10,000 resamplings.<\/p>\n<p>Compared to its predecessor, the Llama 3 8B-based &#8220;DiarizationLM-8b-Fisher-v2&#8221;, this model reduces the error rate to 2.99% on Fisher and 4.92% on Callhome despite having half the parameter count. It also achieves scores significantly lower than the baseline on multi-speaker meeting corpora (ICSI and AMI), which were difficult to evaluate with previous models.<\/p>\n<p>Additionally, benchmark results for the base model &#8220;google\/gemma-4-E4B&#8221; itself have been published as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th style=\"text-align: left;\">Benchmark<\/th>\n<th style=\"text-align: left;\">Gemma 4 31B<\/th>\n<th style=\"text-align: left;\">Gemma 4 26B A4B<\/th>\n<th style=\"text-align: left;\">Gemma 4 12B Unified<\/th>\n<th style=\"text-align: left;\">Gemma 4 E4B<\/th>\n<th style=\"text-align: left;\">Gemma 4 E2B<\/th>\n<th style=\"text-align: left;\">Gemma 3 27B (no think)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"text-align: left;\">MMLU Pro<\/td>\n<td style=\"text-align: left;\">85.2%<\/td>\n<td style=\"text-align: left;\">82.6%<\/td>\n<td style=\"text-align: left;\">77.2%<\/td>\n<td style=\"text-align: left;\">69.4%<\/td>\n<td style=\"text-align: left;\">60.0%<\/td>\n<td style=\"text-align: left;\">67.6%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">AIME 2026 no tools<\/td>\n<td style=\"text-align: left;\">89.2%<\/td>\n<td style=\"text-align: left;\">88.3%<\/td>\n<td style=\"text-align: left;\">77.5%<\/td>\n<td style=\"text-align: left;\">42.5%<\/td>\n<td style=\"text-align: left;\">37.5%<\/td>\n<td style=\"text-align: left;\">20.8%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">LiveCodeBench v6<\/td>\n<td style=\"text-align: left;\">80.0%<\/td>\n<td style=\"text-align: left;\">77.1%<\/td>\n<td style=\"text-align: left;\">72.0%<\/td>\n<td style=\"text-align: left;\">52.0%<\/td>\n<td style=\"text-align: left;\">44.0%<\/td>\n<td style=\"text-align: left;\">29.1%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Codeforces ELO<\/td>\n<td style=\"text-align: left;\">2150<\/td>\n<td style=\"text-align: left;\">1718<\/td>\n<td style=\"text-align: left;\">1659<\/td>\n<td style=\"text-align: left;\">940<\/td>\n<td style=\"text-align: left;\">633<\/td>\n<td style=\"text-align: left;\">110<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">GPQA Diamond<\/td>\n<td style=\"text-align: left;\">84.3%<\/td>\n<td style=\"text-align: left;\">82.3%<\/td>\n<td style=\"text-align: left;\">78.8%<\/td>\n<td style=\"text-align: left;\">58.6%<\/td>\n<td style=\"text-align: left;\">43.4%<\/td>\n<td style=\"text-align: left;\">42.4%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Tau2 (average over 3)<\/td>\n<td style=\"text-align: left;\">76.9%<\/td>\n<td style=\"text-align: left;\">68.2%<\/td>\n<td style=\"text-align: left;\">69.0%<\/td>\n<td style=\"text-align: left;\">42.2%<\/td>\n<td style=\"text-align: left;\">24.5%<\/td>\n<td style=\"text-align: left;\">16.2%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">BigBench Extra Hard<\/td>\n<td style=\"text-align: left;\">74.4%<\/td>\n<td style=\"text-align: left;\">64.8%<\/td>\n<td style=\"text-align: left;\">53.0%<\/td>\n<td style=\"text-align: left;\">33.1%<\/td>\n<td style=\"text-align: left;\">21.9%<\/td>\n<td style=\"text-align: left;\">19.3%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">MMMLU<\/td>\n<td style=\"text-align: left;\">88.4%<\/td>\n<td style=\"text-align: left;\">86.3%<\/td>\n<td style=\"text-align: left;\">83.4%<\/td>\n<td style=\"text-align: left;\">76.6%<\/td>\n<td style=\"text-align: left;\">67.4%<\/td>\n<td style=\"text-align: left;\">70.7%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">MMMU Pro<\/td>\n<td style=\"text-align: left;\">76.9%<\/td>\n<td style=\"text-align: left;\">73.8%<\/td>\n<td style=\"text-align: left;\">69.1%<\/td>\n<td style=\"text-align: left;\">52.6%<\/td>\n<td style=\"text-align: left;\">44.2%<\/td>\n<td style=\"text-align: left;\">49.7%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">MATH-Vision<\/td>\n<td style=\"text-align: left;\">85.6%<\/td>\n<td style=\"text-align: left;\">82.4%<\/td>\n<td style=\"text-align: left;\">79.7%<\/td>\n<td style=\"text-align: left;\">59.5%<\/td>\n<td style=\"text-align: left;\">52.4%<\/td>\n<td style=\"text-align: left;\">46.0%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">CoVoST<\/td>\n<td style=\"text-align: left;\">&#8211;<\/td>\n<td style=\"text-align: left;\">&#8211;<\/td>\n<td style=\"text-align: left;\">38.5*<\/td>\n<td style=\"text-align: left;\">35.54<\/td>\n<td style=\"text-align: left;\">33.47<\/td>\n<td style=\"text-align: left;\">&#8211;<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">FLEURS (lower is better)<\/td>\n<td style=\"text-align: left;\">&#8211;<\/td>\n<td style=\"text-align: left;\">&#8211;<\/td>\n<td style=\"text-align: left;\">0.069*<\/td>\n<td style=\"text-align: left;\">0.08<\/td>\n<td style=\"text-align: left;\">0.09<\/td>\n<td style=\"text-align: left;\">&#8211;<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The base model Gemma 4 E4B records 69.4% on MMLU Pro and 58.6% on GPQA Diamond while being a lightweight model, possessing reasoning performance that surpasses the older generation large model Gemma 3 27B (no think). On the other hand, compared to higher-tier models such as the 31B and 26B A4B, it falls behind in advanced reasoning tasks such as difficult mathematics (AIME 2026) and competitive programming (Codeforces).<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>This model excels at post-processing tasks that correct speaker turn boundary judgments and labeling errors with high precision for text output by ASR and existing speaker diarization systems.<\/p>\n<p>Specifically, while it can accurately correct short backchannels of about 1 to 5 words and lexical turn-taking boundaries, it is designed to maintain acoustic speaker anchors during longer utterances (monologues) of 6 or more words, preventing drift phenomena where speaker labels switch midway. This makes it suitable for utilization in dialog and meeting audio processing such as the following:<\/p>\n<ul>\n<li><strong>Telephone Conversation Text Optimization<\/strong>: Accurately organizing turn overlaps and backchannels from two-speaker calls (such as Fisher) to casual phone conversations involving roughly 2 to 5 people (such as Callhome).<\/li>\n<li><strong>Multi-Participant Meeting Minutes Generation<\/strong>: Suitable for minutes generation and speaker identification in complex environments where multiple people speak actively, such as four-person face-to-face meetings (AMI) or academic research meetings with up to 9 participants (ICSI).<\/li>\n<li><strong>Accuracy Improvement of Existing ASR Pipelines<\/strong>: Improving the Word Diarization Error Rate (WDER) simply by adding format conversion of output text and lightweight inference, without needing to retrain existing speech recognition models or diarization processes.<\/li>\n<\/ul>\n<p>Note that while the base model &#8220;google\/gemma-4-E4B&#8221; is a multimodal foundational model supporting native image and audio processing, this derivative model &#8220;DiarizationLM-Gemma-4-E4B-v1&#8221; is fine-tuned specifically for text processing tasks of taking text-input ASR hypotheses and generating optimized text with speaker labels.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 8.0B parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>8GB (RTX 4060 \/ 3060 Ti, etc.)<\/td>\n<td>Q4_K_M<\/td>\n<td>4.9GB<\/td>\n<td>5.9GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-10-05): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Runs in Ollama, LM Studio and llama.cpp as-is.<\/strong><\/p>\n<p>It is distributed in GGUF, so no conversion is needed.<\/p>\n<p><strong>License \u2014 <code>apache-2.0<\/code> (Commercial use allowed):<\/strong> Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.<\/p>\n<p><strong>Compression:<\/strong> the Q4_K_M build measures 5.26 bits per weight \u2014 about 33% the size of the original 16-bit weights, calculated by this site from the actual file sizes.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:files --><\/p>\n<h2>Distributed Files<\/h2>\n<p><em>Weight files published in <a href=\"https:\/\/huggingface.co\/google\/DiarizationLM-Gemma-4-E4B-v1\/tree\/main\">google\/DiarizationLM-Gemma-4-E4B-v1<\/a>, listed by this site from the Hugging Face API. Sizes are the actual file sizes.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>File<\/th>\n<th>Size<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>DiarizationLM-Gemma-4-E4B-v1-q4_0.gguf<\/code><\/td>\n<td>5.15GB<\/td>\n<\/tr>\n<tr>\n<td><code>DiarizationLM-Gemma-4-E4B-v1-q4_k_m.gguf<\/code><\/td>\n<td>5.30GB<\/td>\n<\/tr>\n<tr>\n<td><code>model.safetensors<\/code><\/td>\n<td>15.99GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:files --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model is available in the Hugging Face repository &#8220;google\/DiarizationLM-Gemma-4-E4B-v1&#8221;. Since it is not a gated model, it can be directly downloaded and used without waiting for additional usage requests or prior approvals.<\/p>\n<p>Weight files are provided in the following formats:<\/p>\n<ul>\n<li><strong>safetensors<\/strong>: A 16-bit (<code>bfloat16<\/code>) format that can be loaded directly with Hugging Face&#8217;s <code>transformers<\/code>.<\/li>\n<li><strong>GGUF<\/strong>: The recommended 4-bit K-quant Medium (<code>Q4_K_M<\/code>) and legacy 4-bit (<code>Q4_0<\/code>) formats are available. The recommended <code>Q4_K_M<\/code> version adopts 256-element superblocks while keeping sensitive layers like <code>attn_v<\/code>, <code>ffn_down<\/code>, and embedding matrices in 6-bit (<code>Q6_K<\/code>).<\/li>\n<\/ul>\n<h3>Execution Steps in Python (transformers + diarizationlm)<\/h3>\n<p>When using in a Python environment, perform GPU inference after installing the related libraries. By combining it with the dedicated library <code>diarizationlm<\/code>, post-processing to reflect LLM output results back into the original hypothesis text can be smoothly executed.<\/p>\n<p>First, install the necessary packages.<\/p>\n<pre><code class=\"language-bash\">pip install transformers diarizationlm\n<\/code><\/pre>\n<p>Next, run model loading and inference with the following code.<\/p>\n<pre><code class=\"language-python\">from diarizationlm import utils\nimport torch\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\n\nMODEL_ID = &quot;google\/DiarizationLM-Gemma-4-E4B-v1&quot;\n\nHYPOTHESIS = (\n    &quot;&lt;speaker:1&gt; Hello, how are you doing &lt;speaker:2&gt; today? I am doing well.&quot;\n    &quot; What about &lt;speaker:1&gt; you? I'm doing well, too. Thank you.&quot;\n)\n\ntokenizer = AutoTokenizer.from_pretrained(MODEL_ID, device_map=&quot;cuda&quot;)\nmodel = AutoModelForCausalLM.from_pretrained(\n    MODEL_ID, torch_dtype=torch.bfloat16, device_map=&quot;cuda&quot;\n)\n\ninputs = tokenizer([HYPOTHESIS + &quot; --&gt; &quot;], return_tensors=&quot;pt&quot;).to(&quot;cuda&quot;)\n\noutputs = model.generate(\n    **inputs,\n    max_new_tokens=int(inputs.input_ids.shape[1] * 1.2),\n    do_sample=False,\n    use_cache=True,\n)\n\ncompletion = tokenizer.batch_decode(\n    outputs[:, inputs.input_ids.shape[1]:], skip_special_tokens=True\n)[0]\ncompletion = utils.truncate_suffix_and_tailing_text(completion, &quot; [eod]&quot;)\n\ntransferred_completion = utils.transfer_llm_completion(completion, HYPOTHESIS)\n\nprint(&quot;Hypothesis:&quot;, HYPOTHESIS)\nprint(&quot;Transferred completion:&quot;, transferred_completion)\n<\/code><\/pre>\n<h3>Execution Steps in llama.cpp<\/h3>\n<p>It is also possible to run the GGUF version directly using inference engines like <code>llama.cpp<\/code>, Ollama, or <code>llama-cpp-python<\/code>. When using <code>llama-cli<\/code>, execute the command as follows:<\/p>\n<pre><code class=\"language-bash\">llama-cli \\\n  -m DiarizationLM-Gemma-4-E4B-v1-q4_k_m.gguf \\\n  -p &quot;&lt;speaker:1&gt; Hello, how are you doing &lt;speaker:2&gt; today? I am doing well. What about &lt;speaker:1&gt; you? I'm doing well, too. Thank you. --&gt; &quot; \\\n  --temp 0.0 \\\n  -n 128\n<\/code><\/pre>\n<p><!-- lmw:same-task --><\/p>\n<h2>Other Models for the Same Task<\/h2>\n<p><em>Recent vision-language models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site&#8217;s estimate.<\/em><\/p>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/05\/qwen-3-8-flash-next-released\/\">Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory<\/a> (180.0B, Over 80GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/04\/glm-5-3-flash-gguf-released\/\">GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory<\/a> (321.3B, Over 80GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/clef-flash-specs-performance-and-use-cases\/\">clef-flash Vision-Language Model: Our Test Answers, 4GB+ VRAM<\/a> (9.4B, 4GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/ggml-org-openjev-gguf-released-2\/\">OpenJev-GGUF 27.4B Decision-Making Specialized Model: 24GB+ VRAM<\/a> (27.4B, 24GB)<\/li>\n<\/ul>\n<p><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-vision\">See all vision-language models \u2192<\/a><\/p>\n<p><!-- \/lmw:same-task --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model runs from the 8GB tier) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-8gb-en\/\">Other models that run on a 8GB GPU<\/a><\/li>\n<li><strong>Engines that run this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<li><strong>What Q4_K_M mean and where to get this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">Quantization and model-format glossary<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-google-en\/\">Google: models, licenses and articles<\/a><\/li>\n<li><strong>How to read MMLU-Pro, AIME, LiveCodeBench<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-benchmarks-en\/\">Benchmark glossary<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/google\/DiarizationLM-Gemma-4-E4B-v1\">https:\/\/huggingface.co\/google\/DiarizationLM-Gemma-4-E4B-v1<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/google\/gemma-4-E4B\">https:\/\/huggingface.co\/google\/gemma-4-E4B<\/a><\/li>\n<li><a href=\"https:\/\/arxiv.org\/abs\/2401.03506\">https:\/\/arxiv.org\/abs\/2401.03506<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/google\/speaker-id\/tree\/master\/DiarizationLM\">https:\/\/github.com\/google\/speaker-id\/tree\/master\/DiarizationLM<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Learn about DiarizationLM-Gemma-4-E4B-v1, an open-weights LLM for speech processing by Google based on Gemma 4 E4B, with specs, performance, and usage.<\/p>\n","protected":false},"author":1,"featured_media":9816,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[1032,2982,163,2984,2986,759,1547,1952,2988],"class_list":["post-9817","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-gemma-en","tag-gemma-4-e4b-en","tag-gguf-en","tag-google-en","tag-google-diarizationlm-gemma-4-e4b-v1-en","tag-transformers-en","tag-verified","tag-vlm-en","tag--en"],"lang":"en","translations":{"en":9817,"ja":9815},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9817","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=9817"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9817\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/9816"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=9817"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=9817"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=9817"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}