{"id":6423,"date":"2026-09-28T13:50:34","date_gmt":"2026-09-28T04:50:34","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/orcasaq2-27b-qwen3-8-27b-quantized\/"},"modified":"2026-09-28T19:23:52","modified_gmt":"2026-09-28T10:23:52","slug":"orcasaq2-27b-qwen3-8-27b-quantized","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/orcasaq2-27b-qwen3-8-27b-quantized\/","title":{"rendered":"OrcaSAQ-2-27B Text Generation Model: 16GB+ VRAM"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/orcarouter\/OrcaSAQ-2-27B\">orcarouter\/OrcaSAQ-2-27B<\/a><\/td>\n<\/tr>\n<tr>\n<td>Family guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-qwen-qwen3-8-en\/\">Qwen3.8 guide (8 articles)<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-alibaba-en\/\">Alibaba (Qwen): models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-24<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-exl-en\/\">EXL3<\/a> \/ safetensors<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Unverified (not confirmed by a primary source)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>OrcaRouter, an AI gateway operator, has released a quantized version of Qwen3.8-27B compressed to an average of 3.21 bits, named <strong>OrcaSAQ2 27B<\/strong> (<code>orcarouter\/OrcaSAQ-2-27B<\/code>), on Hugging Face. The original model, which is 54GB in BF16, has been reduced to a total of 12.3GB across its distribution files. The stated goal of the publisher is to &#8220;<strong>run a 27B model on a single 16GB GPU<\/strong>,&#8221; and they present numerical comparisons directly against the BF16 version to show that the difference from the original model is small.<\/p>\n<p>The publisher calls this a &#8220;proprietary sensitivity-aware mixed-precision quantization system (SAQ2)&#8221; and has not disclosed the details of the method (calibration procedures, bit-width allocation per layer, or packing methods). However, looking at the distribution files&#8217; configuration (<code>quantization_config.json<\/code>), the quantization format is <code>exl3<\/code> (ExLlamaV3 format, version 1.5.1). The publisher&#8217;s concurrently released execution kernel repository also describes its contents as &#8220;a combination of EXL3 trellis quantization (QTIP family) with mixed-precision allocation determined by search.&#8221; In other words, it is accurate to read this as <strong>not a new format, but rather a clever allocation of bit-widths per layer within the existing EXL3 format<\/strong>.<\/p>\n<p>The base Qwen3.8-27B model itself is covered in <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B Multimodal Vision-Language Model: 8GB+ VRAM, GGUF Builds<\/a>.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Base Model: Qwen3.8-27B (Hybrid architecture of the Qwen3.5 family. Out of 64 layers, 48 are Gated DeltaNet and 16 are standard attention)<\/li>\n<li>Quantization: EXL3 format, averaging 3.21 bits for the decoder part. Output layer (<code>lm_head<\/code>) is 6-bit, embeddings are int8, and MTP heads are 4-bit<\/li>\n<li>Distribution Files: Split into 4 safetensors files, total 12.27GB<\/li>\n<li>Context Length: 262,144 tokens (design upper limit. The length actually usable on a 16GB GPU is described later)<\/li>\n<li>Supported Features: Thinking mode, tool calling, speculative decoding via MTP (Multi-Token Prediction)<\/li>\n<li>Input: Text only. While the base Qwen3.8-27B can read images and videos, this version does not include a visual encoder<\/li>\n<li>License: Apache-2.0 (inherited from the base Qwen3.8-27B)<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<h3>Deviation from the Original Model (BF16)<\/h3>\n<p>The publisher provides values comparing the distributed weights themselves using the same evaluation pipeline as the BF16 version, measured on 16,376 tokens of WikiText-2.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Build<\/th>\n<th>Size<\/th>\n<th>Decoder Bit-width<\/th>\n<th>Mean KLD \u2193<\/th>\n<th>Top-1 Match Rate \u2191<\/th>\n<th>PPL \u2193<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Qwen3.8-27B BF16<\/td>\n<td>54GB<\/td>\n<td>16<\/td>\n<td>\u2014<\/td>\n<td>100%<\/td>\n<td>5.6468<\/td>\n<\/tr>\n<tr>\n<td>OrcaSAQ2 27B<\/td>\n<td>12.3GB<\/td>\n<td>3.21<\/td>\n<td>0.031<\/td>\n<td>93.2%<\/td>\n<td>5.6482<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The meanings of the three metrics are as follows:<\/p>\n<ul>\n<li><strong>PPL (Perplexity)<\/strong>: Measures how well the model can predict the continuation of text without confusion. Lower is better. The degradation from BF16 is +0.02%, showing almost no difference<\/li>\n<li><strong>KLD (Kullback-Leibler Divergence)<\/strong>: Measures how much the probability distribution of the next token deviates from BF16. Closer to 0 means more faithful to the original model<\/li>\n<li><strong>Top-1 Match Rate<\/strong>: The proportion of times the &#8220;most likely next token chosen&#8221; matched the BF16 version<\/li>\n<\/ul>\n<p>It is important to read these three metrics side by side. Even though the PPL difference is only 0.02%, the Top-1 match rate is 93.2%, meaning <strong>roughly once every 15 tokens, a different token is chosen as the most likely candidate compared to BF16<\/strong>. Because PPL is an averaged value across the entire text, individual choices tend to cancel each other out and do not readily show up in the aggregate score. The publisher itself notes in the limitations section that &#8220;a 93.2% match rate means some token decisions differ from BF16&#8221; and &#8220;a +0.02% PPL does not guarantee identical performance on downstream tasks.&#8221; Rather than summarizing it as &#8220;virtually lossless,&#8221; it is more accurate to understand it as &#8220;<strong>average predictions barely change, but individual decisions change at a certain rate<\/strong>.&#8221; Note that comparisons measuring other quantized versions (like GGUF or other EXL3 variants) under the same conditions are not provided, so it is impossible to judge from this data alone whether a KLD of 0.031 is exceptional for the 3-bit range.<\/p>\n<h3>Agent Benchmarks (Publisher&#8217;s Self-Report)<\/h3>\n<p>The model card presents scores of 70.0 on SWE-bench Verified and 58.4 on Terminal-Bench 2.1 in a table alongside published values for Claude Sonnet 4.6, Gemini 3, GPT-5.4, and others. However, <strong>this article does not treat these as rankings against other models<\/strong> for the following reasons:<\/p>\n<ul>\n<li>The publisher itself clarifies that the table is &#8220;not a controlled comparison, but published reference values.&#8221; Agent benchmark scores vary significantly depending on the execution environment (harness), thinking budget, and time limits used<\/li>\n<li>The specific execution environment and settings used to measure the OrcaSAQ2 scores are not documented<\/li>\n<li>The most crucial piece of information\u2014&#8221;<strong>the score of the BF16 version measured under the exact same conditions<\/strong>&#8220;\u2014is missing. Therefore, it is impossible to tell from this table how much performance was lost due to quantization<\/li>\n<\/ul>\n<p>The only meaningful evaluation for the quantized version is the &#8220;deviation from BF16&#8221; discussed in the previous section. It is best not to read more into the agent benchmark scores than simply that the publisher reported them.<\/p>\n<h3>Speed on a 16GB GPU<\/h3>\n<p>Values measured with vLLM while restricting GPU memory to 15.7GiB.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Setting<\/th>\n<th>1 Stream<\/th>\n<th>8 Streams<\/th>\n<th>16 Streams<\/th>\n<th>KV Cache Capacity<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Without MTP<\/td>\n<td>65.3 tok\/s<\/td>\n<td>332 tok\/s<\/td>\n<td>333 tok\/s<\/td>\n<td>29,354 tokens<\/td>\n<\/tr>\n<tr>\n<td>With MTP<\/td>\n<td>90.1 tok\/s<\/td>\n<td>220 tok\/s<\/td>\n<td>219 tok\/s<\/td>\n<td>14,563 tokens<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Using MTP (speculative decoding) makes generation about 38% faster for single-user use. In exchange, because the draft head takes KV cache from the same memory pool, <strong>the available context length is cut roughly in half, and it actually becomes slower when handling many concurrent requests<\/strong>. The choice comes down to using MTP enabled for single-user interactive use, and disabled for batching multiple requests.<\/p>\n<p>The execution kernel repository contains a table with the same intent, but the numbers differ (with a 15.5GiB limit, 1 stream with MTP is 111.2 tok\/s, and KV cache without MTP is 59,753 tokens). This is likely due to differences in settings or versions, but both are measurements by the publisher, and independent third-party replication is not yet available.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>The target use case envisioned by the publisher is agents performing long procedures involving coding, terminal operations, and browser navigation. The publisher&#8217;s concern is that even small errors from quantization can cause a single tool call to change, cascading to alter all subsequent states\u2014which is why tasks with long procedures are more susceptible to quantization effects, and why they detailed the &#8220;deviation from BF16&#8221; so thoroughly.<\/p>\n<p>From a local user&#8217;s perspective, this model is suited for those who want to <strong>set up a 27B-class model as a server using a 16GB GPU (such as an RTX 5080 or 4080)<\/strong>. Since it runs as an OpenAI-compatible API, it can be integrated with coding agents or custom tools.<\/p>\n<p>There are also many points to keep in mind:<\/p>\n<ul>\n<li><strong>You cannot use a 262K context on 16GB.<\/strong> Because the weights consume about 12.3GB, only about 3 to 4GB remains. The publisher also recommends &#8220;starting with a context around 32K on a 16GB GPU and adjusting according to the use case&#8221;<\/li>\n<li><strong>Images cannot be read.<\/strong> The visual capabilities of the base Qwen3.8-27B have been stripped out in this version<\/li>\n<li><strong>The quantization method is private.<\/strong> You cannot apply the exact same procedure to other models yourself<\/li>\n<li>The publisher is a company operating an AI gateway, and many models released on Hugging Face are uncensored versions of existing models. This OrcaSAQ2 27B is released as a direct quantization of Qwen3.8-27B rather than an uncensored version<\/li>\n<\/ul>\n<h2>How It Differs from Similar Models<\/h2>\n<ul>\n<li><strong>Difference from general GGUF quantized versions (Q3_K_M, IQ3 families, etc.)<\/strong>: While the file sizes are similar, the execution engines are completely different. While GGUF runs out of the box in llama.cpp, Ollama, and LM Studio, this version does not run in them (see the next section). Its strengths lie in providing numerical measurements of deviation from BF16 and the ability to handle multiple concurrent requests using vLLM<\/li>\n<li><strong>Difference from other EXL3 quantized versions<\/strong>: The EXL3 format itself is the standard format for ExLlamaV3, and other versions around 3 bits could potentially be created. The differences in OrcaSAQ2 lie in searching for layer-wise bit allocation distributions and providing a plugin to run it with vLLM. However, numerical comparisons against other EXL3 versions under identical conditions are not provided<\/li>\n<li><strong>Difference from the base Qwen3.8-27B (BF16)<\/strong>: The required memory drops to under a quarter, but image input capability is lost, and token-level agreement is limited to 93.2%. The choice is between the original model if accuracy is the top priority, or this version if fitting into 16GB is the priority<\/li>\n<\/ul>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 27.8B parameters (taken from the base model Qwen\/Qwen3.8-27B)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>16GB (RTX 5060 Ti 16GB \/ 4060 Ti 16GB, etc.)<\/td>\n<td>EXL3<\/td>\n<td>11.4GB<\/td>\n<td>13.7GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Not usable in Ollama, LM Studio and llama.cpp yet \u2014 we have found no GGUF build.<\/strong><\/p>\n<p>The publisher ships EXL3 \/ safetensors only. Today it can be run with transformers, using the memory figures in the table above.<\/p>\n<p><strong>License \u2014 <code>apache-2.0<\/code> (Commercial use allowed):<\/strong> Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<p>Values we measured ourselves on our server (no GPU) by actually reading and running this model&#8217;s files \u2014 not figures copied from the model card. How we measure, and the results for every model: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h3>Japanese Token Efficiency<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Tokenizer<\/th>\n<th>Tokens per 1,000 Japanese characters<\/th>\n<th>Ratio to the same text in English<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>This model<\/strong><\/td>\n<td><strong>546<\/strong><\/td>\n<td><strong>0.99\u00d7<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Qwen3<\/td>\n<td>688<\/td>\n<td>1.26\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Llama 3.2<\/td>\n<td>744<\/td>\n<td>1.36\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Gemma 3<\/td>\n<td>564<\/td>\n<td>1.03\u00d7<\/td>\n<\/tr>\n<tr>\n<td>gpt-oss<\/td>\n<td>795<\/td>\n<td>1.45\u00d7<\/td>\n<\/tr>\n<tr>\n<td>LLM-jp-3<\/td>\n<td>497<\/td>\n<td>0.85\u00d7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>It needs about 21% fewer tokens than the Qwen3 tokenizer for the same Japanese text, so about 1.26\u00d7 as much Japanese fits in the same context length, and generation is faster per character.<\/p>\n<p>Counted with the <code>tokenizer.json<\/code> of <a href=\"https:\/\/huggingface.co\/orcarouter\/OrcaSAQ-2-27B\">orcarouter\/OrcaSAQ-2-27B<\/a> on a fixed text we wrote ourselves (876 Japanese characters across news, conversation, technical docs, a formal email, travel writing and a recipe) and its English translation. Fewer tokens mean more Japanese fits in the context window.<\/p>\n<p><!-- \/lmw:lab --><\/p>\n<p><!-- lmw:peers --><\/p>\n<h2>Recent Models in the Same Size Class<\/h2>\n<p><em>Models with <\/em><em>15\u201340B<\/em><em> parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site&#8217;s estimates; licenses are as stated on the model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Our article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>XingChen-AGI\/Xing4.0-29B-A4B<\/td>\n<td>31.2B<\/td>\n<td>24GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/xing4-0-29b-a4b-china-telecom\/\">Xing4.0-29B-A4B 29B MoE Model Strong in Coding Agents: 24GB+ VRAM<\/a> (2026-09-28)<\/td>\n<\/tr>\n<tr>\n<td>Edge0\/Edge0-35B-A3B-preview<\/td>\n<td>36.0B<\/td>\n<td>24GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/edge0-35b-a3b-preview-sparse-moe\/\">Edge0-35B-A3B-preview 35B MoE Model for Phone-Class Memory: 24GB+ VRAM<\/a> (2026-09-11)<\/td>\n<\/tr>\n<tr>\n<td>bartowski\/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF<\/td>\n<td>26.5B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/pantheon-reasoning-26b-gguf\/\">Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF: 12GB+ VRAM<\/a> (2026-09-11)<\/td>\n<\/tr>\n<tr>\n<td>nex-agi\/Nex-N2.5-mini<\/td>\n<td>35.1B<\/td>\n<td>16GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-n25-mini-released\/\">Nex-N2.5-mini Agent Model for Long-Horizon Tasks: 16GB+ VRAM<\/a> (2026-09-09)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:peers --><\/p>\n<h2>How to Get It<\/h2>\n<ul>\n<li>Distribution: Hugging Face <code>orcarouter\/OrcaSAQ-2-27B<\/code> (safetensors, EXL3 format)<\/li>\n<li><strong>Does not run in llama.cpp, Ollama, or LM Studio.<\/strong> The execution kernel repository explicitly states: &#8220;Since llama.cpp can only read its own quantization codebooks, these weights cannot be used. Converting them to GGUF would require re-quantizing with a different quantizer&#8221;<\/li>\n<li><strong>When running with vLLM<\/strong>: In addition to vLLM, you need to install the publisher&#8217;s plugin (<code>Continuum-AI-Corp\/OrcaSAQ2-kernel<\/code>). The model card suggests <code>pip install git+https:\/\/github.com\/Continuum-AI-Corp\/OrcaSAQ2-kernel<\/code>, and the kernel repository also provides startup configurations and Docker images tailored for 16GB. To correctly receive thinking outputs and tool calls, append <code>--reasoning-parser qwen3<\/code> and <code>--tool-call-parser qwen3_xml<\/code><\/li>\n<li><strong>When running with ExLlamaV3<\/strong>: The kernel repository notes that &#8220;for single-user use on consumer GPUs, it is lighter than vLLM and better suited.&#8221; However, you must first apply a patch to read embeddings packed in int8. Note that using MTP in ExLlamaV3 allocates separate caching for draft heads, increasing memory usage by about 10GB, which means it will not fit on a 16GB GPU<\/li>\n<li>Recommended sampling settings: temperature 1.0, top_p 0.95, top_k 20. Thinking mode is enabled by default<\/li>\n<li>The execution kernel repository was just created on September 24, making it very recent since release. It is wise to anticipate potential issues where it may not run smoothly depending on your environment<\/li>\n<li>License: Model is Apache-2.0 (inherited from the base Qwen3.8-27B). Please review the original license text before use<\/li>\n<\/ul>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/hemmingway-1-open-27b-model-specialized-for-human-like-writing\/\">Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/\">Signal-3.8-27B-GGUF Token-Efficient Optimized GGUF Model: 16GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B Multimodal Vision-Language Model: 8GB+ VRAM, GGUF Builds<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/13\/qwen3-8-27b-twin-turbo-fable-cold-fusion-709-l-gguf\/\">Qwen3.8-27B TWIN-TURBO Uncensored GGUF Released<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model runs from the 16GB tier) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-16gb-en\/\">Other models that run on a 16GB GPU<\/a><\/li>\n<li><strong>Explore the same model family<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-qwen-qwen3-8-en\/\">Qwen3.8 family overview (8 articles, 11 converted builds)<\/a><\/li>\n<li><strong>What EXL3 mean and where to get this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-exl-en\/\">EXL2 \/ EXL3 format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">Quantization and model-format glossary<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-alibaba-en\/\">Alibaba (Qwen): models, licenses and articles<\/a><\/li>\n<li><strong>Other models for the same task<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-text\">Other text generation models<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/orcarouter\/OrcaSAQ-2-27B\">https:\/\/huggingface.co\/orcarouter\/OrcaSAQ-2-27B<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/Continuum-AI-Corp\/OrcaSAQ2-kernel\">https:\/\/github.com\/Continuum-AI-Corp\/OrcaSAQ2-kernel<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3.8-27B\">https:\/\/huggingface.co\/Qwen\/Qwen3.8-27B<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-09-28: Added our own measurements: Japanese token efficiency.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n<blockquote>\n<p><strong>This article contains unverified information.<\/strong> We will append an update note once it is confirmed by a primary source.<\/p>\n<\/blockquote>\n","protected":false},"excerpt":{"rendered":"<p>OrcaRouter has released OrcaSAQ2 27B, a compressed 3.21-bit quantization of Qwen3.8-27B designed to run on a single 16GB GPU.<\/p>\n","protected":false},"author":1,"featured_media":6422,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[2592,2594,2596,522,1082,1565,592,1555],"class_list":["post-6423","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-exl3-en","tag-orcarouter-en","tag-orcarouter-orcasaq-2-27b-en","tag-qwen3-8-en","tag-qwen3-8-27b-en","tag-unverified","tag-vllm-en","tag--en"],"lang":"en","translations":{"en":6423,"ja":6421},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6423","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=6423"}],"version-history":[{"count":3,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6423\/revisions"}],"predecessor-version":[{"id":6803,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6423\/revisions\/6803"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/6422"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=6423"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=6423"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=6423"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}