{"id":2388,"date":"2026-09-21T22:26:50","date_gmt":"2026-09-21T13:26:50","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/model-qwen-qwen-image-2-1-en\/"},"modified":"2026-09-28T05:32:20","modified_gmt":"2026-09-27T20:32:20","slug":"model-qwen-qwen-image-2-1-en","status":"publish","type":"page","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-qwen-qwen-image-2-1-en\/","title":{"rendered":"Qwen-Image-2.1 Guide: Model Files"},"content":{"rendered":"<h2>About This Model<\/h2>\n<p><strong>Qwen-Image-2.1<\/strong> is an image generation model from Alibaba&#8217;s Qwen team. Its distinguishing feature is that <strong>text-to-image generation and image editing live in one model<\/strong>, and its image-generation component is compact for this class of model, at <strong>7B parameters<\/strong> (32 DiT layers). It outputs images of about 4 megapixels, such as 2048\u00d72048 at 1:1.<\/p>\n<p>The publisher highlights four improvements:<\/p>\n<ul>\n<li><strong>Native support for transparent (RGBA) images.<\/strong> The same model can generate images with transparent backgrounds from text, edit transparent layers, and cut subjects out of photos. For logos, icons and compositing assets, no separate background-removal step is needed.<\/li>\n<li><strong>Up to 10 reference images.<\/strong> You can place a person or product in a new scene while preserving their appearance. Areas to edit can be marked with circles, painted annotations, or a separate mask.<\/li>\n<li><strong>Better typography, portrait lighting and fine texture<\/strong>, according to the publisher.<\/li>\n<li>A lightweight design (mixed-granularity attention and prefix KV cache reuse) keeps compute costs down.<\/li>\n<\/ul>\n<p><strong>About the &#8220;PE&#8221; companion models:<\/strong> Two helper models that rewrite short requests into detailed image prompts were released alongside it. <strong>PE-T2I<\/strong> (for text-to-image) turns a brief request in any language into a detailed English prompt plus a recommended aspect ratio; <strong>PE-I2I<\/strong> (for image editing) takes a vague editing instruction and the input image(s) and produces a precise, actionable instruction. Both are fine-tuned from Qwen3.5-VL 9B. <strong>This site&#8217;s article on the official models and the &#8220;Distributed Files&#8221; table below cover PE-I2I, not the image generator itself.<\/strong><\/p>\n<h2>What Makes It Stand Out<\/h2>\n<p>The model card does not include a numeric comparison with other models, so this page does not discuss benchmark numbers and focuses on capabilities instead.<\/p>\n<ul>\n<li><strong>Generation, editing and transparency handled by one model<\/strong> is the headline. Work that used to require a generation model, an editing model and a background-removal tool can run in a single pipeline.<\/li>\n<li><strong>The small 7B generation component<\/strong> helps when running on local GPUs, though total memory use\u2014including the text-understanding part\u2014depends on the format you use.<\/li>\n<\/ul>\n<h2>Running It Locally<\/h2>\n<ul>\n<li>The official route is diffusers&#8217; <code>QwenImage21Pipeline<\/code> (at the time of writing, you need to install diffusers from GitHub).<\/li>\n<li>Community GGUF builds are also available. Articles on this site include an &#8220;Uncensored&#8221; build from a third party with the safety tuning removed; its weights differ from the official model, so it is a separate model.<\/li>\n<li><strong>The license is the Qwen Research License, which limits use to research and evaluation.<\/strong> Commercial use requires a separate license from the publisher, and redistributions must include the required attribution notice.<\/li>\n<\/ul>\n<p><em>Sources: model cards and license for <a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1\">Qwen\/Qwen-Image-2.1<\/a>, <a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1-PE-T2I\">Qwen\/Qwen-Image-2.1-PE-T2I<\/a> and <a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1-PE-I2I\">Qwen\/Qwen-Image-2.1-PE-I2I<\/a>, as of 2026-09-25.<\/em><\/p>\n<h2>Our Coverage and Data<\/h2>\n<p>Everything Local Model Watch has published about the <strong>Qwen-Image-2.1<\/strong> family: 3 article(s) covering the base model and its fine-tunes, plus converted builds we tracked after publication. Memory requirements below are computed by this site from file sizes, not quoted from model cards. Part of our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-en\/\">model family index<\/a>.<\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Base model(s)<\/td>\n<td><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1\">Qwen\/Qwen-Image-2.1<\/a>, <a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1-PE-T2I\">Qwen\/Qwen-Image-2.1-PE-T2I<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-alibaba-en\/\">Alibaba (Qwen)<\/a><\/td>\n<\/tr>\n<tr>\n<td>License (model card)<\/td>\n<td>other<\/td>\n<\/tr>\n<tr>\n<td>Articles<\/td>\n<td>3<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Can You Run It Locally?<\/h2>\n<p>The publisher distributes this model as safetensors.<\/p>\n<p><strong>License \u2014 <code>other<\/code>:<\/strong> A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats and the license field. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><em>This assessment is for Qwen\/Qwen-Image-2.1-PE-I2I.<\/em><\/p>\n<h2>Distributed Files<\/h2>\n<p><em>Weight files published in <a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1-PE-I2I\/tree\/main\">Qwen\/Qwen-Image-2.1-PE-I2I<\/a>, listed by this site from the Hugging Face API. Sizes are the actual file sizes.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>File<\/th>\n<th>Size<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>model-00001.safetensors<\/code><\/td>\n<td>5.07GB<\/td>\n<\/tr>\n<tr>\n<td><code>model-00002.safetensors<\/code><\/td>\n<td>5.09GB<\/td>\n<\/tr>\n<tr>\n<td><code>model-00003.safetensors<\/code><\/td>\n<td>5.09GB<\/td>\n<\/tr>\n<tr>\n<td><code>model-00004.safetensors<\/code><\/td>\n<td>3.57GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Quantized and Converted Variants<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Added<\/th>\n<th>Publisher<\/th>\n<th>Format<\/th>\n<th>Repository<\/th>\n<th>Smallest VRAM tier (build, est. memory)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-23<\/td>\n<td>unsloth<\/td>\n<td>GGUF<\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/Qwen-Image-2.1-GGUF\">unsloth\/Qwen-Image-2.1-GGUF<\/a><\/td>\n<td>Q3_K_XL 4.0GB (fits in 4GB VRAM)<\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td>unsloth<\/td>\n<td>FP8<\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/Qwen-Image-2.1-FP8\">unsloth\/Qwen-Image-2.1-FP8<\/a><\/td>\n<td>FP8 8.0GB (fits in 8GB VRAM)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>File sizes of each build:<\/p>\n<ul>\n<li>Available builds in unsloth\/Qwen-Image-2.1-GGUF: Q2_K 2.3GB \/ Q3_K_S 2.5GB \/ Q3_K_M 3.0GB \/ Q3_K_XL 3.4GB \/ Q4_K_S 3.6GB \/ Q4_K_M 3.9GB \/ Q5_K_S 4.2GB \/ Q5_K_M 5.0GB \/ Q6_K 5.8GB \/ Q6_K_XL 6.3GB \/ Q8_0 7.1GB \/ F16 13.3GB<\/li>\n<li>Available builds in unsloth\/Qwen-Image-2.1-FP8: FP8 6.6GB \/ INT8 6.8GB \/ FP8(Qwen-Image-2.1-text_encoder-FP8) 8.7GB<\/li>\n<\/ul>\n<p><em>This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files.<\/em><\/p>\n<h2>Articles (the family&#8217;s own models first, then newest)<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Published<\/th>\n<th>Model<\/th>\n<th>Type<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-20<\/td>\n<td>Qwen\/Qwen-Image-2.1-PE-I2I<\/td>\n<td>Image, Video and Audio<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/20\/qwen-image-2-1-released\/\">Qwen-Image-2.1 Released: Open-Weight Image Gen &amp; Editing<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td>Viggle\/Qwen-Image-2.1-viggle-turbo<\/td>\n<td>Image, Video and Audio<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/qwen-image-2-1-viggle-turbo-v0-1-preview\/\">Qwen-Image-2.1-viggle-turbo Image Generation Model: 48GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-21<\/td>\n<td>abenzerps\/Qwen-Image-2.1-Uncensored-GGUF<\/td>\n<td>Image, Video and Audio<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/uncensored-qwen-image-2-1-gguf-released\/\">Qwen-Image-2.1-Uncensored-GGUF Image Generation Model: 16GB+ VRAM<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Repositories<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1\">Qwen\/Qwen-Image-2.1<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1-PE-T2I\">Qwen\/Qwen-Image-2.1-PE-T2I<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Viggle\/Qwen-Image-2.1-viggle-turbo\">Viggle\/Qwen-Image-2.1-viggle-turbo<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/abenzerps\/Qwen-Image-2.1-Uncensored-GGUF\">abenzerps\/Qwen-Image-2.1-Uncensored-GGUF<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen-Image-2.1-PE-I2I\">Qwen\/Qwen-Image-2.1-PE-I2I<\/a><\/li>\n<\/ul>\n<p><em>Last updated 2026-09-28 (JST). The explanation at the top of this page was written with the help of AI from the primary sources it cites. The tables and lists under &#8220;Our Coverage and Data&#8221; are assembled by code from our article log.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>About This Model Qwen-Image-2.1 is an image generation model from Alibaba&#8217;s Qwen team. Its distinguishing feature is that text-to-image generation and image editing live in one model, and its image-generation component is compact for this class of model, at 7B parameters (32 DiT layers). It outputs images of about 4 megapixels, such as 2048\u00d72048 at [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-2388","page","type-page","status-publish","hentry"],"lang":"en","translations":{"en":2388,"ja":2387},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/2388","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=2388"}],"version-history":[{"count":14,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/2388\/revisions"}],"predecessor-version":[{"id":5992,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/2388\/revisions\/5992"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=2388"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}