{"id":575,"date":"2026-09-12T21:16:35","date_gmt":"2026-09-12T12:16:35","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/"},"modified":"2026-09-18T21:42:00","modified_gmt":"2026-09-18T12:42:00","slug":"signal-3-8-27b-gguf-overview","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/","title":{"rendered":"Signal-3.8-27B-GGUF: Faster and More Token-Efficient"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/agentionai\/Signal-3.8-27B-GGUF\">agentionai\/Signal-3.8-27B-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-10<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>GGUF<\/td>\n<\/tr>\n<tr>\n<td>Paper<\/td>\n<td><a href=\"https:\/\/arxiv.org\/abs\/2606.00206\">arXiv:2606.00206<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p><code>agentionai\/Signal-3.8-27B-GGUF<\/code> is a GGUF model created by minimally invasively fine-tuning the base model Qwen3.8-27B to operate with lower generation latency and higher token efficiency. In general prompt evaluations, it successfully reduces response tokens by 57% and reasoning tokens by 52% while maintaining answer quality equal to or better than the base model. This allows it to complete processing in less than half the wall time of the base model on equivalent hardware.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>License: apache-2.0<\/li>\n<li>Parameters: 27.8B<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>According to the model card evaluations, compared to the base model Qwen3.8-27B Q8_0, Signal significantly reduces response and reasoning tokens. Below are the measured results of token changes in general prompts and coding.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th><\/th>\n<th>base Q8_0<\/th>\n<th>Signal<\/th>\n<th>change<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>general answers, median tokens<\/td>\n<td>243<\/td>\n<td>104<\/td>\n<td><strong>-57%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>answers opening with a preamble (&#8220;Sure!&#8221;, &#8220;Great question&#8221;)<\/td>\n<td>13%<\/td>\n<td>0%<\/td>\n<td><strong>gone<\/strong><\/td>\n<\/tr>\n<tr>\n<td>answers with markdown headers<\/td>\n<td>47%<\/td>\n<td>18%<\/td>\n<td>-62%<\/td>\n<\/tr>\n<tr>\n<td>answers with bold<\/td>\n<td>85%<\/td>\n<td>52%<\/td>\n<td>-39%<\/td>\n<\/tr>\n<tr>\n<td>coding answers, median tokens<\/td>\n<td>159<\/td>\n<td>142<\/td>\n<td>-11%<\/td>\n<\/tr>\n<tr>\n<td>coding answers, p90 tokens<\/td>\n<td>1026<\/td>\n<td>914<\/td>\n<td>-11%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Token counts during the thinking mode were also measured.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th><\/th>\n<th>base Q8_0<\/th>\n<th>Signal<\/th>\n<th>change<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>reasoning tokens, general prompts, median<\/td>\n<td>153<\/td>\n<td>74<\/td>\n<td><strong>-52%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>reasoning tokens, coding prompts, median<\/td>\n<td>225<\/td>\n<td>166<\/td>\n<td>-26%<\/td>\n<\/tr>\n<tr>\n<td>reasoning tokens, GSM8K, median<\/td>\n<td>119<\/td>\n<td>81<\/td>\n<td>-32%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Exact match quality on GSM8K (a benchmark measuring step-by-step arithmetic word problem solving at the elementary school level) is as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th><\/th>\n<th>base Q8_0<\/th>\n<th>Signal<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>thinking off, 60 problems<\/td>\n<td>98.3%<\/td>\n<td>98.3%<\/td>\n<\/tr>\n<tr>\n<td>thinking on, 40 problems<\/td>\n<td>92.5%<\/td>\n<td><strong>95.0%<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Additionally, speed comparisons using speculative decoding are as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>prompt \/ draft length<\/th>\n<th style=\"text-align: right;\">base acceptance<\/th>\n<th style=\"text-align: right;\">Signal acceptance<\/th>\n<th style=\"text-align: right;\">decode speed vs base<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>prose, draft 3<\/td>\n<td style=\"text-align: right;\">39%<\/td>\n<td style=\"text-align: right;\">47%<\/td>\n<td style=\"text-align: right;\"><strong>+10%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>prose, draft 4<\/td>\n<td style=\"text-align: right;\">35%<\/td>\n<td style=\"text-align: right;\">28%<\/td>\n<td style=\"text-align: right;\">-9%<\/td>\n<\/tr>\n<tr>\n<td>structured output (JSON), draft 3<\/td>\n<td style=\"text-align: right;\">72%<\/td>\n<td style=\"text-align: right;\">94%<\/td>\n<td style=\"text-align: right;\"><strong>+20%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>structured output (JSON), draft 4<\/td>\n<td style=\"text-align: right;\">66%<\/td>\n<td style=\"text-align: right;\">87%<\/td>\n<td style=\"text-align: right;\"><strong>+22%<\/strong><\/td>\n<\/tr>\n<tr>\n<td>chat prompts, sampled at 0.7, adaptive draft \u22644 (40 prompts)<\/td>\n<td style=\"text-align: right;\">57%<\/td>\n<td style=\"text-align: right;\">60%<\/td>\n<td style=\"text-align: right;\">\u2014<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The performance of the original model Qwen3.8-27B has also been measured across various benchmarks. A partial excerpt of the text performance comparison table is as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th><\/th>\n<th>Qwen3.8-27B<\/th>\n<th>Qwen3.6-27B<\/th>\n<th>Qwen3.7-Plus<\/th>\n<th>Muse Glimmer-30B<\/th>\n<th>Opus4.6 Max<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Coding<\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>Agentic terminal coding Terminal Bench 2.1 (Terminus)<\/td>\n<td>73.0<\/td>\n<td>63.4<\/td>\n<td>64.0<\/td>\n<td>51.7<\/td>\n<td>78.2<\/td>\n<\/tr>\n<tr>\n<td>Agentic coding SWE-bench Pro<\/td>\n<td>61.7<\/td>\n<td>53.5<\/td>\n<td>57.6<\/td>\n<td>51.2<\/td>\n<td>53.4<\/td>\n<\/tr>\n<tr>\n<td>Repo-level code generation NL2Repo-Bench<\/td>\n<td>42.3<\/td>\n<td>36.2<\/td>\n<td>41.1<\/td>\n<td>&#8212;<\/td>\n<td>47.6<\/td>\n<\/tr>\n<tr>\n<td>Agentic coding DeepSWE 1.1<\/td>\n<td>42.2<\/td>\n<td>13.3<\/td>\n<td>14.2<\/td>\n<td>&#8212;<\/td>\n<td>&#8212;<\/td>\n<\/tr>\n<tr>\n<td>Software engineering QwenSWEBench<\/td>\n<td>79.0<\/td>\n<td>49.3<\/td>\n<td>59.2<\/td>\n<td>&#8212;<\/td>\n<td>63.8<\/td>\n<\/tr>\n<tr>\n<td>Agent<\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>Long-horizon office work CoWorkBench<\/td>\n<td>70.7<\/td>\n<td>61.0<\/td>\n<td>65.1<\/td>\n<td>&#8212;<\/td>\n<td>68.2<\/td>\n<\/tr>\n<tr>\n<td>Professional job tasks JobBench<\/td>\n<td>33.4<\/td>\n<td>21.8<\/td>\n<td>27.6<\/td>\n<td>&#8212;<\/td>\n<td>&#8212;<\/td>\n<\/tr>\n<tr>\n<td>Frontier agentic tasks Agents&#8217; Last Exam<\/td>\n<td>Pass@1 20.4 Score 42.9<\/td>\n<td>Pass@1 10.6 Score 27.3<\/td>\n<td>Pass@1 13.2 Score 33.6<\/td>\n<td>&#8212;<\/td>\n<td>&#8212;<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>Instruction following IFBench<\/td>\n<td>79.5<\/td>\n<td>69.1<\/td>\n<td>79.1<\/td>\n<td>77.0<\/td>\n<td>62.5<\/td>\n<\/tr>\n<tr>\n<td>Scientific reasoning GPQA Diamond<\/td>\n<td>89.2<\/td>\n<td>87.8<\/td>\n<td>90.3<\/td>\n<td>83.5<\/td>\n<td>91.3<\/td>\n<\/tr>\n<tr>\n<td>Multidisciplinary reasoning HLE<\/td>\n<td>30.8<\/td>\n<td>24.0<\/td>\n<td>34.7<\/td>\n<td>22.0<\/td>\n<td>40.0<\/td>\n<\/tr>\n<tr>\n<td>Competitive coding LiveCodeBench v6<\/td>\n<td>90.3<\/td>\n<td>83.9<\/td>\n<td>89.6<\/td>\n<td>&#8212;<\/td>\n<td>88.8<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>These figures show that while maintaining the base model&#8217;s excellent knowledge, reasoning capabilities, and high potential in code generation and agent tasks, Signal achieves significant speedups and improved token efficiency by stripping away unnecessary preambles and redundant reasoning steps.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>Signal is a fine-tuned model of Qwen3.8-27B designed to achieve lower generation latency and higher token efficiency. It cuts down on token usage without sacrificing answer quality by providing direct answers and omitting unnecessary preambles, excessive formatting, greetings, and explanatory narration. Furthermore, even in thinking mode, it reduces the tokens used to explain the process while preserving useful reasoning steps. It also retains the image input capabilities (vision encoder and projector) inherited from the base model.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 27.8B parameters (taken from the base model Qwen\/Qwen3.8-27B)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>16GB (RTX 5060 Ti 16GB \/ 4060 Ti 16GB, etc.)<\/td>\n<td>IQ4_XS<\/td>\n<td>13.3GB<\/td>\n<td>15.9GB<\/td>\n<\/tr>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>Q5_K_M<\/td>\n<td>18.2GB<\/td>\n<td>21.8GB<\/td>\n<\/tr>\n<tr>\n<td>32GB (RTX 5090, etc.)<\/td>\n<td>Q6_K<\/td>\n<td>20.9GB<\/td>\n<td>25.1GB<\/td>\n<\/tr>\n<tr>\n<td>48GB (RTX 6000 Ada \/ A6000, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>27.1GB<\/td>\n<td>32.5GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<h2>How to Get It<\/h2>\n<ul>\n<li>Distribution format: GGUF<\/li>\n<li>Supported engines: llama.cpp (llama-server, etc.), Ollama, LM Studio<\/li>\n<\/ul>\n<p>Example command for running with llama.cpp:<\/p>\n<pre><code class=\"language-bash\">llama-server -hf agentionai\/Signal-3.8-27B-GGUF:AP-Q4_K_XL \\\n  --jinja -ngl 999 -fa on -c 32768 \\\n  --temp 0.7 --top-p 0.95 --top-k 20 --min-p 0\n<\/code><\/pre>\n<p>To add the built-in draft head to increase throughput, append the following arguments (requires a build supporting <code>--spec-type draft-mtp<\/code>):<\/p>\n<pre><code class=\"language-bash\">  --spec-type draft-mtp --spec-draft-n-max 4\n<\/code><\/pre>\n<p>Note that when using image inputs, download and use <code>mmproj-BF16.gguf<\/code> (0.87 GiB) for the base model located at the root of the repository together with any tier.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">Salience-27B-R6 GGUF Released by bartowski<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/agentionai\/Signal-3.8-27B-GGUF\">https:\/\/huggingface.co\/agentionai\/Signal-3.8-27B-GGUF<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3.8-27B\">https:\/\/huggingface.co\/Qwen\/Qwen3.8-27B<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Discover Signal-3.8-27B-GGUF, a minimally invasive fine-tune of Qwen3.8-27B offering lower latency, reduced token usage, and high performance.<\/p>\n","protected":false},"author":1,"featured_media":574,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[1053,163,1073,1082,1084,117],"class_list":["post-575","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-fine-tuning-en","tag-gguf-en","tag-llama-cpp-en","tag-qwen3-8-27b-en","tag-signal-en","tag--en"],"lang":"en","translations":{"en":575,"ja":573},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/575","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=575"}],"version-history":[{"count":8,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/575\/revisions"}],"predecessor-version":[{"id":1718,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/575\/revisions\/1718"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/574"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=575"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=575"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=575"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}