{"id":627,"date":"2026-09-14T20:34:33","date_gmt":"2026-09-14T11:34:33","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf-quantizations\/"},"modified":"2026-09-18T21:42:01","modified_gmt":"2026-09-18T12:42:01","slug":"orion-26b-a4b-v1-1-gguf-quantizations","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf-quantizations\/","title":{"rendered":"Orion-26B-A4B-v1.1 GGUF Quants Released by bartowski"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/TheDrummer_Orion-26B-A4B-v1.1-GGUF\">bartowski\/TheDrummer_Orion-26B-A4B-v1.1-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-14<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>GGUF<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>It is reported that bartowski has released GGUF quantization files for the multimodal model Orion-26B-A4B-v1.1. As some information included could not be officially verified at the time of writing, caution is advised when deploying or testing the model in practice. It supports text, image, and audio inputs, and a variety of quantization files are provided for local execution.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Parameters: 26B<\/li>\n<li>Input Support: Text, image, audio (requires mmproj file)<\/li>\n<li>Speculative Decoding: None<\/li>\n<li>imatrix: Supported<\/li>\n<li>Architecture: Gemma4ForConditionalGeneration<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>Regarding the performance of the base model, according to the model card, a low temperature setting (such as 0.9) is optimal, and Gemma&#8217;s thinking capabilities are reported to be very sharp in roleplay (RP). Below is a comparison table of per-tensor layouts for the published quantized versions.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Quant<\/th>\n<th>Size<\/th>\n<th>Body bits\/weight<\/th>\n<th>File bits\/weight<\/th>\n<th>Body kept at base type<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Q6_K_L<\/td>\n<td>24.52GB<\/td>\n<td>7.38<\/td>\n<td>7.60<\/td>\n<td>51 %<\/td>\n<\/tr>\n<tr>\n<td>Q6_K<\/td>\n<td>23.79GB<\/td>\n<td>7.03<\/td>\n<td>7.37<\/td>\n<td>71 %<\/td>\n<\/tr>\n<tr>\n<td>Q6_K_S<\/td>\n<td>23.14GB<\/td>\n<td>6.71<\/td>\n<td>7.17<\/td>\n<td>90 %<\/td>\n<\/tr>\n<tr>\n<td>Q5_K_M<\/td>\n<td>20.01GB<\/td>\n<td>5.92<\/td>\n<td>6.20<\/td>\n<td>71 %<\/td>\n<\/tr>\n<tr>\n<td>Q5_K_S<\/td>\n<td>18.40GB<\/td>\n<td>5.70<\/td>\n<td>5.70<\/td>\n<td>90 %<\/td>\n<\/tr>\n<tr>\n<td>Q4_K_L<\/td>\n<td>18.21GB<\/td>\n<td>5.19<\/td>\n<td>5.64<\/td>\n<td>50 %<\/td>\n<\/tr>\n<tr>\n<td>Q4_K_M<\/td>\n<td>17.97GB<\/td>\n<td>5.07<\/td>\n<td>5.57<\/td>\n<td>70 %<\/td>\n<\/tr>\n<tr>\n<td>Q4_K_S<\/td>\n<td>16.05GB<\/td>\n<td>4.70<\/td>\n<td>4.98<\/td>\n<td>90 %<\/td>\n<\/tr>\n<tr>\n<td>IQ4_NL<\/td>\n<td>15.63GB<\/td>\n<td>5.07<\/td>\n<td>4.85<\/td>\n<td>70 %<\/td>\n<\/tr>\n<tr>\n<td>Q3_K_L<\/td>\n<td>14.80GB<\/td>\n<td>4.39<\/td>\n<td>4.59<\/td>\n<td>50 %<\/td>\n<\/tr>\n<tr>\n<td>IQ3_M<\/td>\n<td>14.34GB<\/td>\n<td>4.17<\/td>\n<td>4.45<\/td>\n<td>50 %<\/td>\n<\/tr>\n<tr>\n<td>IQ4_XS<\/td>\n<td>14.28GB<\/td>\n<td>4.42<\/td>\n<td>4.43<\/td>\n<td>90 %<\/td>\n<\/tr>\n<tr>\n<td>Q3_K_M<\/td>\n<td>13.79GB<\/td>\n<td>4.01<\/td>\n<td>4.27<\/td>\n<td>70 %<\/td>\n<\/tr>\n<tr>\n<td>IQ3_XS<\/td>\n<td>12.84GB<\/td>\n<td>3.57<\/td>\n<td>3.98<\/td>\n<td>90 %<\/td>\n<\/tr>\n<tr>\n<td>IQ3_XXS<\/td>\n<td>12.61GB<\/td>\n<td>3.46<\/td>\n<td>3.91<\/td>\n<td>70 %<\/td>\n<\/tr>\n<tr>\n<td>Q3_K_S<\/td>\n<td>12.37GB<\/td>\n<td>3.56<\/td>\n<td>3.84<\/td>\n<td>90 %<\/td>\n<\/tr>\n<tr>\n<td>Q2_K<\/td>\n<td>11.61GB<\/td>\n<td>3.19<\/td>\n<td>3.60<\/td>\n<td>70 %<\/td>\n<\/tr>\n<tr>\n<td>IQ2_M<\/td>\n<td>10.90GB<\/td>\n<td>2.89<\/td>\n<td>3.38<\/td>\n<td>70 %<\/td>\n<\/tr>\n<tr>\n<td>IQ2_S<\/td>\n<td>10.31GB<\/td>\n<td>2.61<\/td>\n<td>3.20<\/td>\n<td>70 %<\/td>\n<\/tr>\n<tr>\n<td>IQ2_XS<\/td>\n<td>9.90GB<\/td>\n<td>2.41<\/td>\n<td>3.07<\/td>\n<td>90 %<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>(Table truncated. Subsequent rows omitted)<\/p>\n<p>Furthermore, the comparative verification results between the computed layout and the standard layout are shown below.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Quant<\/th>\n<th>Computed layout KLD<\/th>\n<th>Standard layout KLD<\/th>\n<th>Ratio at equal size<\/th>\n<th>Size vs standard file<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Q4_K_M<\/td>\n<td>0.8086 \u00b1 0.0113<\/td>\n<td>1.1462 \u00b1 0.0132<\/td>\n<td>0.88\u00d7<\/td>\n<td>+5.5 %<\/td>\n<\/tr>\n<tr>\n<td>Q3_K_M<\/td>\n<td>1.2642 \u00b1 0.0149<\/td>\n<td>2.9859 \u00b1 0.0241<\/td>\n<td>0.51\u00d7<\/td>\n<td>+5.9 %<\/td>\n<\/tr>\n<tr>\n<td>IQ2_XS<\/td>\n<td>2.6730 \u00b1 0.0206<\/td>\n<td>2.9355 \u00b1 0.0210<\/td>\n<td>0.91\u00d7<\/td>\n<td>\u22122.3 %<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>From these tables, it is confirmed that quantization files adopting the computed layout, such as Q4_K_M and Q3_K_M, keep the Kullback-Leibler divergence (KLD) lower compared to the standard layout, resulting in smaller differences from the unquantized model and maintained quality. Q4_K_M in particular shows a favorable balance.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>This model supports image and audio input in addition to text, and is intended for conversational use and roleplay (RP) applications. Inheriting the characteristics of the base model, it is well-suited for interactive tasks leveraging multimodal inputs.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 25.8B parameters (taken from the base model TheDrummer\/Orion-26B-A4B-v1.1)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>8GB (RTX 4060 \/ 3060 Ti, etc.)<\/td>\n<td>F16<\/td>\n<td>1.1GB<\/td>\n<td>1.3GB<\/td>\n<\/tr>\n<tr>\n<td>12GB (RTX 4070 \/ 3060 12GB, etc.)<\/td>\n<td>IQ2_S<\/td>\n<td>9.6GB<\/td>\n<td>11.5GB<\/td>\n<\/tr>\n<tr>\n<td>16GB (RTX 5060 Ti 16GB \/ 4060 Ti 16GB, etc.)<\/td>\n<td>IQ3_M<\/td>\n<td>13.4GB<\/td>\n<td>16.0GB<\/td>\n<\/tr>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>Q5_K_M<\/td>\n<td>18.6GB<\/td>\n<td>22.4GB<\/td>\n<\/tr>\n<tr>\n<td>32GB (RTX 5090, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>25.0GB<\/td>\n<td>30.0GB<\/td>\n<\/tr>\n<tr>\n<td>80GB class (A100 \/ H100)<\/td>\n<td>BF16<\/td>\n<td>48.1GB<\/td>\n<td>57.8GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<h2>How to Get It<\/h2>\n<ul>\n<li>Distribution Format: GGUF<\/li>\n<li>Supported Engines: llama.cpp, LM Studio, koboldcpp, ramalama, Jan AI, Text Generation Web UI, LoLLMs, Atomic Chat<\/li>\n<li>Example Download Command:<\/li>\n<\/ul>\n<pre><code>hf download bartowski\/TheDrummer_Orion-26B-A4B-v1.1-GGUF --include &quot;TheDrummer_Orion-26B-A4B-v1.1-Q4_K_M.gguf&quot; --local-dir.\/\n<\/code><\/pre>\n<p>When utilizing multimodal inputs, you will also need to prepare the corresponding multimodal projector file (<code>mmproj-TheDrummer_Orion-26B-A4B-v1.1-f16.gguf<\/code> or <code>mmproj-TheDrummer_Orion-26B-A4B-v1.1-bf16.gguf<\/code>).<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/intern-s2-397b-gguf-released\/\">Intern-S2-397B GGUF Quantized Models Released by bartowski<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/bartowski-nex-agi-nex-n2-5-mini-gguf-2\/\">bartowski\/nex-agi_Nex-N2.5-mini-GGUF: Specs and Hardware<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">Salience-27B-R6 GGUF Released by bartowski<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/pantheon-reasoning-26b-gguf\/\">Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/bartowski\/TheDrummer_Orion-26B-A4B-v1.1-GGUF\">https:\/\/huggingface.co\/bartowski\/TheDrummer_Orion-26B-A4B-v1.1-GGUF<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/TheDrummer\/Orion-26B-A4B-v1.1\">https:\/\/huggingface.co\/TheDrummer\/Orion-26B-A4B-v1.1<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\">https:\/\/github.com\/ggml-org\/llama.cpp<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/bartowski1182\/quantization-config\">https:\/\/github.com\/bartowski1182\/quantization-config<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/blog\/bartowski\/per-tensor-layout-maps-for-gguf-quantization\">https:\/\/huggingface.co\/blog\/bartowski\/per-tensor-layout-maps-for-gguf-quantization<\/a><\/li>\n<li><a href=\"https:\/\/gist.github.com\/bartowski1182\/e26453c0404e24eb317543ec5360f87a\">https:\/\/gist.github.com\/bartowski1182\/e26453c0404e24eb317543ec5360f87a<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/9921\">https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/9921<\/a><\/li>\n<li><a href=\"https:\/\/gist.github.com\/Artefact2\/b5f810600771265fc1e39442288e8ec9\">https:\/\/gist.github.com\/Artefact2\/b5f810600771265fc1e39442288e8ec9<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Explore bartowski&#8217;s GGUF quantizations for Orion-26B-A4B-v1.1, featuring multimodal support, performance tables, and hardware requirements.<\/p>\n","protected":false},"author":1,"featured_media":626,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[1030,163,1073,1178,896,117],"class_list":["post-627","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-bartowski-en","tag-gguf-en","tag-llama-cpp-en","tag-orion-26b-a4b-v1-1-en","tag--en"],"lang":"en","translations":{"en":627,"ja":625},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/627","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=627"}],"version-history":[{"count":7,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/627\/revisions"}],"predecessor-version":[{"id":1664,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/627\/revisions\/1664"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/626"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=627"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=627"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=627"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}