{"id":672,"date":"2026-09-15T20:16:56","date_gmt":"2026-09-15T11:16:56","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/15\/salience-27b-r6-gguf-2\/"},"modified":"2026-09-18T21:42:03","modified_gmt":"2026-09-18T12:42:03","slug":"salience-27b-r6-gguf-2","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/","title":{"rendered":"Salience-27B-R6 GGUF Released by bartowski"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/vectionlabs_Salience-27B-R6-GGUF\">bartowski\/vectionlabs_Salience-27B-R6-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-14<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>GGUF<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<p>bartowski has released the GGUF quantized version of &#8220;Salience-27B-R6&#8221;, developed by Vection Labs, as <code>bartowski\/vectionlabs_Salience-27B-R6-GGUF<\/code>.<\/p>\n<p>This model is a full-parameter active 27B Dense model based on Qwen3.8, supporting text as well as image and video inputs. It implements &#8220;Reasoning economy&#8221; at the weight level to suppress excessive consumption of thinking tokens while maintaining thinking capabilities, aiming for improved efficiency particularly in long-term agent loop processing.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Parameters: 27.8B (Dense)<\/li>\n<li>Context Length: 1,048,576 tokens (YaRN + Dual Chunk Attention)<\/li>\n<li>License: apache-2.0<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>According to the original model card, measurement results from official benchmark suites have not been published (unmeasured).<\/p>\n<p>On the other hand, as part of the model&#8217;s design policy, optimizations have been implemented to intentionally shorten the reasoning chain required to reach the same conclusion compared to the previous version (R5). This is reported to significantly reduce the overall task completion time in agent workloads involving multi-step tool calls. However, it is explained that as a trade-off for shortening the reasoning, scores on multiple-choice knowledge benchmarks and the like may drop slightly compared to the previous version.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<ul>\n<li>Software engineering and coding agents (code generation, debugging, repository-wide changes, etc.).<\/li>\n<li>Terminal operation and tool integration (multi-step command planning, execution result verification, error recovery).<\/li>\n<li>Technical analysis leveraging multimodal inputs (reading architecture diagrams, UI screenshots, stack trace images, whiteboard photos, etc.).<\/li>\n<li>Built-in MTP (Multi-Token Prediction) head, enabling fast inference via self-speculative decoding in compatible environments.<\/li>\n<\/ul>\n<h2>Differences from Similar Models<\/h2>\n<p>Other articles covering derivatives of the same base model (Qwen3.8-27B) include <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/12\/signal-3-8-27b-gguf\/\">Qwen3.8-27B Fast Fine-tune &#8220;Signal-3.8-27B-GGUF&#8221; Released<\/a>, which aims to reduce token count and improve response speed, and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/06\/qwopus3-8-27b-flash-gguf\/\">Qwopus3.8-27B-Flash GGUF Released, Lightweight Model for Agents<\/a>, which features lightweight optimizations for agents.<\/p>\n<p>In contrast, Salience-27B-R6 differs by adopting a 27B Dense configuration where all parameters operate constantly rather than an MoE configuration, building short-thinking optimization directly into the model weights rather than through configuration changes, and providing native support for multimodal inputs (images and video) and ultra-long contexts exceeding 1 million tokens.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 27.8B parameters (taken from the base model vectionlabs\/Salience-27B-R6)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>12GB (RTX 4070 \/ 3060 12GB, etc.)<\/td>\n<td>IQ2_M<\/td>\n<td>9.8GB<\/td>\n<td>11.8GB<\/td>\n<\/tr>\n<tr>\n<td>16GB (RTX 5060 Ti 16GB \/ 4060 Ti 16GB, etc.)<\/td>\n<td>Q3_K_L<\/td>\n<td>13.2GB<\/td>\n<td>15.8GB<\/td>\n<\/tr>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>Q5_K_M<\/td>\n<td>19.5GB<\/td>\n<td>23.4GB<\/td>\n<\/tr>\n<tr>\n<td>32GB (RTX 5090, etc.)<\/td>\n<td>Q6_K_L<\/td>\n<td>23.2GB<\/td>\n<td>27.9GB<\/td>\n<\/tr>\n<tr>\n<td>48GB (RTX 6000 Ada \/ A6000, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>27.1GB<\/td>\n<td>32.5GB<\/td>\n<\/tr>\n<tr>\n<td>80GB class (A100 \/ H100)<\/td>\n<td>BF16<\/td>\n<td>50.9GB<\/td>\n<td>61.1GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<h2>How to Get It<\/h2>\n<p>It is distributed in GGUF format and can be downloaded from Hugging Face.<\/p>\n<p>In a llama.cpp environment, you can install and start the server with the following commands:<\/p>\n<pre><code class=\"language-bash\">curl -LsSf https:\/\/llama.app\/install.sh | sh\nllama-server -hf bartowski\/vectionlabs_Salience-27B-R6-GGUF:Q4_K_M\n<\/code><\/pre>\n<p>Example command for individual download using the Hugging Face CLI:<\/p>\n<pre><code class=\"language-bash\">hf download bartowski\/vectionlabs_Salience-27B-R6-GGUF --include &quot;vectionlabs_Salience-27B-R6-Q4_K_M.gguf&quot; --local-dir.\/\n<\/code><\/pre>\n<p>In addition to llama.cpp, this model is compatible with LM Studio, koboldcpp, ramalama, Jan AI, Text Generation Web UI, LoLLMs, Atomic Chat, and others.<\/p>\n<p>When using image recognition, you need to specify (<code>--mmproj<\/code>) the bundled multimodal projector file (<code>mmproj-vectionlabs_Salience-27B-R6-f16.gguf<\/code> or <code>bf16<\/code>). Also, to use acceleration via MTP, append <code>--spec-type draft-mtp<\/code> to the command line.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/\">Signal-3.8-27B-GGUF: Faster and More Token-Efficient<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/pantheon-reasoning-26b-gguf\/\">Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/intern-s2-397b-gguf-released\/\">Intern-S2-397B GGUF Quantized Models Released by bartowski<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf-quantizations\/\">Orion-26B-A4B-v1.1 GGUF Quants Released by bartowski<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/bartowski\/vectionlabs_Salience-27B-R6-GGUF\">https:\/\/huggingface.co\/bartowski\/vectionlabs_Salience-27B-R6-GGUF<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/vectionlabs\/Salience-27B-R6\">https:\/\/huggingface.co\/vectionlabs\/Salience-27B-R6<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>bartowski has released GGUF quantizations for Vection Labs&#8217; Salience-27B-R6, a multimodal 27B Dense model with reasoning economy.<\/p>\n","protected":false},"author":1,"featured_media":671,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[1030,163,136,1257,1259,896,117],"class_list":["post-672","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-bartowski-en","tag-gguf-en","tag-llm-en","tag-salience-27b-r6-en","tag-vection-labs-en","tag--en"],"lang":"en","translations":{"en":672,"ja":670},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/672","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=672"}],"version-history":[{"count":7,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/672\/revisions"}],"predecessor-version":[{"id":1674,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/672\/revisions\/1674"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/671"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=672"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=672"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=672"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}