{"id":2384,"date":"2026-09-21T22:26:45","date_gmt":"2026-09-21T13:26:45","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/21\/qwen-image-2-1-gguf-released\/"},"modified":"2026-09-22T10:20:54","modified_gmt":"2026-09-22T01:20:54","slug":"qwen-image-2-1-gguf-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/qwen-image-2-1-gguf-released\/","title":{"rendered":"Qwen-Image-2.1-GGUF Released: Local Image Generation"},"content":{"rendered":"<p><em>Sample outputs are available on the <a href=\"https:\/\/huggingface.co\/abenzerps\/Qwen-Image-2.1-GGUF\">model card<\/a>.<\/em><\/p>\n<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/abenzerps\/Qwen-Image-2.1-GGUF\">abenzerps\/Qwen-Image-2.1-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-21<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>other<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>GGUF \/ safetensors<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<p><em>Sample outputs are available on the <a href=\"https:\/\/huggingface.co\/abenzerps\/Qwen-Image-2.1-GGUF\">model card<\/a>.<\/em><\/p>\n<h2>Overview<\/h2>\n<p>It is reported that a GGUF quantized version repository for the latest model &#8220;Qwen-Image-2.1&#8221;, which supports text-to-image generation and advanced image editing, has been released as &#8220;<a href=\"https:\/\/huggingface.co\/abenzerps\/Qwen-Image-2.1-GGUF\">abenzerps\/Qwen-Image-2.1-GGUF<\/a>&#8220;. Note that this news is unconfirmed information that has not been officially verified by the developers or other official sources, and is considered an unofficial quantized release by a third party. The model is distributed with the aim of enabling text-to-image generation of regular images and transparent images with alpha channels (RGBA), as well as editing existing images and extracting subjects, directly in local PC environments, according to reports.<\/p>\n<h2>Specifications<\/h2>\n<p>Based on public information and the specifications of the original model &#8220;Qwen-Image-2.1&#8221;, the specs are as follows:<\/p>\n<ul>\n<li><strong>Parameter Count<\/strong>: 7B (7 billion parameters) in the visual generation component<\/li>\n<li><strong>Architecture<\/strong>: 32-layer Single-Stream DiT (employing mixed-granularity attention and prefix KV cache reuse structure)<\/li>\n<li><strong>Output Specifications<\/strong>:<\/li>\n<li>Supports standard image generation and transparent image generation with alpha channels (RGBA) &#8211; Supports editing transparent layers, extracting subjects from photos, and editing using up to 10 reference images &#8211; Supported aspect ratios and resolutions: &#8211; 1:1 (2048 x 2048) &#8211; 4:3 (2400 x 1792) &#8211; 3:4 (1792 x 2400) &#8211; 3:2 (2528 x 1696) &#8211; 2:3 (1696 x 2528) &#8211; 16:9 (2752 x 1536) &#8211; 9:16 (1536 x 2752)<\/li>\n<li><strong>Recommended Settings<\/strong>:<\/li>\n<li>Inference steps: 40 steps (from the original model&#8217;s quick start)<\/li>\n<li><strong>License<\/strong>: Qwen Research License (detailed conditions such as commercial use comply with the license terms)<\/li>\n<\/ul>\n<h2>Performance and Quality<\/h2>\n<p>According to explanations from the developers of the original model &#8220;Qwen-Image-2.1&#8221;, the following four main improvements have been made since the previous generation, which are claimed to enhance quality and efficiency:<\/p>\n<ol>\n<li><strong>Compact and Efficient<\/strong>: Lightweight architectural design is said to provide strong image quality with low computational cost.<\/li>\n<li><strong>Native Transparency Support<\/strong>: Features the ability to integratively process text-to-RGBA transparent image generation, transparent layer editing, and subject extraction from photos within a single model.<\/li>\n<li><strong>Versatile Editing Functions<\/strong>: Supports up to 10 reference images, enabling local editing with circle designations, drawing annotations, and independent masks. It is said to allow editing while maintaining the identity of people or products.<\/li>\n<li><strong>Realistic Textures and Refined Aesthetics<\/strong>: Improvements in typography (text rendering), portrait lighting, and fine details are reported to yield more visually appealing results.<\/li>\n<\/ol>\n<p>Additionally, according to the model card of the newly released GGUF quantized version repository, this release is adjusted to be an &#8220;Uncensored&#8221; specification. It is reported that it lacks built-in safety checkers or content filters, making it possible to directly generate adult content (NSFW) or sensitive images without prompt refusals or output image blackouts. Therefore, the resulting behavior is explained to be completely dependent on the input prompt and the execution environment.<\/p>\n<p>Regarding the operational stability of the GGUF version, it is reported that the Q8_0 quantization (<code>qwen-image-2.1-Q8_0.gguf<\/code>) may cause shape mismatch errors (<code>[136] vs [128]<\/code>) depending on the GPU or ComfyUI environment. Therefore, users seeking stable operation are recommended to use the Q4_K_M, Q5_K_M, or Q6_K quantization versions.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>Based on the specifications of the original model &#8220;Qwen-Image-2.1&#8221;, the use cases where this model excels are said to be diverse.<\/p>\n<p>Specifically, alongside high-quality image generation from text prompts, native generation of transparent images including alpha channels (RGBA) is highlighted. It is also explained to be suited for tasks such as image editing utilizing transparent layers and subject extraction from photos (background transparency processing).<\/p>\n<p>Furthermore, it supports up to 10 reference images, enabling image editing while maintaining the identity (features) of people or products. During editing, it supports area designation via circles, handwritten annotations, and local editing (partial editing) using individual mask images, making it well-suited for pinpoint correction work.<\/p>\n<p>Text rendering (typography) accuracy has also improved, and it is reported to exhibit high quality in adding text to images, portrait lighting, and expressing fine details.<\/p>\n<p>In this GGUF version, because it exhibits &#8220;Uncensored&#8221; behavior where built-in safety filters are not applied, it is said to be suited for use cases where users wish to avoid prompt refusals and output blackouts, directly generating sensitive images including adult content (NSFW).<\/p>\n<h2>How It Differs from Similar Models<\/h2>\n<p>As a related past article, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/21\/abenzerps-qwen-image-21-uncensored-gguf\/\">Image Generation Model &#8220;abenzerps\/Qwen-Image-2.1-Uncensored-GGUF&#8221; Released<\/a> can be cited. Similar to the model introduced in the previous article, the model distributed in the current repository &#8220;<a href=\"https:\/\/huggingface.co\/abenzerps\/Qwen-Image-2.1-GGUF\">abenzerps\/Qwen-Image-2.1-GGUF<\/a>&#8221; is also explained to possess practically uncensored characteristics. The biggest difference in this repository is that not only the GGUF format transformer model, but also the companion files required for operation such as text encoders and VAEs, are all packaged and provided within the same repository. This is said to save users the trouble of separately searching for files from multiple locations and allow for a smoother integration into ComfyUI environments.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 7.1B parameters (taken from the base model Qwen\/Qwen-Image-2.1)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>8GB (RTX 4060 \/ 3060 Ti, etc.)<\/td>\n<td>Q6_K<\/td>\n<td>5.5GB<\/td>\n<td>6.6GB<\/td>\n<\/tr>\n<tr>\n<td>12GB (RTX 4070 \/ 3060 12GB, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>7.1GB<\/td>\n<td>8.5GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model and related files are available from the Hugging Face repository. The distribution formats are GGUF and safetensors.<\/p>\n<p>It is stated that the model can be run in a local environment by combining ComfyUI with the &#8220;ComfyUI-GGUF&#8221; extension.<\/p>\n<h3>File Placement<\/h3>\n<p>Place the downloaded files according to the following directory structure in ComfyUI:<\/p>\n<pre><code class=\"language-text\">ComfyUI\/\n\u2514\u2500\u2500 models\/\n    \u251c\u2500\u2500 diffusion_models\/\n    \u2502   \u2514\u2500\u2500 qwen-image-2.1-Q4_K_M.gguf         # GGUF quantized model (Q4_K_M recommended)\n    \u251c\u2500\u2500 text_encoders\/\n    \u2502   \u2514\u2500\u2500 qwen3vl_8b_bf16.safetensors        # Text encoder (int8 version recommended for memory saving)\n    \u2514\u2500\u2500 vae\/\n        \u2514\u2500\u2500 qwen_image_2.1_vae_bf16.safetensors # VAE file\n<\/code><\/pre>\n<h3>ComfyUI Setup<\/h3>\n<ol>\n<li>\n<p><strong>Update and Install ComfyUI-GGUF<\/strong><br \/>\n   Clone the <code>leejet\/ComfyUI-GGUF<\/code> repository, which natively supports Qwen-Image 2.1, into ComfyUI&#8217;s <code>custom_nodes<\/code> directory. <code>bash cd ComfyUI\/custom_nodes git clone https:\/\/github.com\/leejet\/ComfyUI-GGUF<\/code> <em>Note: If you are using the older <code>city96\/ComfyUI-GGUF<\/code> and encounter an &#8220;Unknown model architecture!&#8221; error, it is stated that you need to update to the aforementioned <code>leejet<\/code> version or add <code>ModelQwenImage<\/code> to the conversion script.<\/em><\/p>\n<\/li>\n<li>\n<p><strong>Node Configuration<\/strong><br \/>\n   &#8211; <strong>Diffusion Model<\/strong>: Add the <code>Unet Loader (GGUF)<\/code> node and select the downloaded GGUF file. &#8211; <strong>Text Encoder<\/strong>: Add the standard <code>CLIPLoader<\/code> node, select the text encoder (<code>qwen3vl_8b_bf16.safetensors<\/code> or <code>int8<\/code> version), and set <code>type<\/code> to <code>qwen_image<\/code>. &#8211; <strong>VAE<\/strong>: Add the standard <code>VAELoader<\/code> node and select <code>qwen_image_2.1_vae_bf16.safetensors<\/code>.<\/p>\n<\/li>\n<li>\n<p><strong>Applying the Workflow<\/strong><br \/>\n   You can use Comfy-Org&#8217;s official workflow templates (for Text-to-Image or Image Edit) as a base. It is explained that you can run it by replacing the default <code>UNETLoader<\/code> node in the workflow with <code>Unet Loader (GGUF)<\/code>.<\/p>\n<\/li>\n<\/ol>\n<h3>Recommended Operating Settings<\/h3>\n<p>The configuration recommended is to keep the GGUF format Diffusion model itself, which is directly tied to sampling speed, in the GPU&#8217;s VRAM, while offloading the text encoder\u2014which runs only once during prompt processing\u2014to system RAM (CPU). It is reported that this significantly reduces VRAM consumption with almost no impact on generation speed. Additionally, if an error occurs due to insufficient VRAM capacity, it is recommended to add the <code>--lowvram<\/code> argument when starting ComfyUI.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/uncensored-qwen-image-2-1-gguf-released\/\">Uncensored Qwen-Image-2.1 GGUF Released for Local ComfyUI<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/abenzerps\/Qwen-Image-2.1-GGUF\">https:\/\/huggingface.co\/abenzerps\/Qwen-Image-2.1-GGUF<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1\">https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-09-22: Verified the content against the official primary source.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>A GGUF quantized version of Qwen-Image-2.1 has been released, supporting text-to-image generation and advanced image editing locally.<\/p>\n","protected":false},"author":1,"featured_media":2383,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[1852,677,163,1835,1547,377,753,1884],"class_list":["post-2384","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-abenzerps-en","tag-comfyui-en","tag-gguf-en","tag-qwen-image-2-1-en","tag-verified","tag--en","tag-ai-en"],"lang":"en","translations":{"en":2384,"ja":2382},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2384","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=2384"}],"version-history":[{"count":1,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2384\/revisions"}],"predecessor-version":[{"id":2452,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2384\/revisions\/2452"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/2383"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=2384"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=2384"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=2384"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}