{"id":7372,"date":"2026-09-29T23:20:01","date_gmt":"2026-09-29T14:20:01","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/29\/longlive-plug-wan21-ti2v-5b-cfg-lora\/"},"modified":"2026-09-30T02:13:47","modified_gmt":"2026-09-29T17:13:47","slug":"longlive-plug-wan21-ti2v-5b-cfg-lora","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/29\/longlive-plug-wan21-ti2v-5b-cfg-lora\/","title":{"rendered":"LongLive-Plug-Wan2.2-TI2V-5B-cfg Video Generation Model: 24GB+ VRAM"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/Efficient-Large-Model\/LongLive-Plug-Wan2.2-TI2V-5B-cfg\">Efficient-Large-Model\/LongLive-Plug-Wan2.2-TI2V-5B-cfg<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-efficient-large-model-en\/\">Efficient Large Model (NVIDIA and MIT): models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-29<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>LongLive-Plug-Wan2.2-TI2V-5B-cfg has been released as a LoRA adapter applying CFG (Classifier-Free Guidance) distillation for the text- and image-to-video generation model Wan2.2-TI2V-5B.<\/p>\n<p>This model is an adapter that distills the standard CFG=5 guided flow of the base model into a single conditional forward pass at CFG=1 on the student model side. It eliminates the need to calculate unconditional (negative) branches at each denoising step, reducing the DiT (Diffusion Transformer) forward processing to just once per step. Note that denoising schedule distillation is not performed, making this different from DMD (Distribution Matching Distillation) checkpoints that carry out few-step generation.<\/p>\n<h2>Specifications<\/h2>\n<p>The specifications and recommended settings listed in the model card are as follows:<\/p>\n<ul>\n<li>Parameter count: LoRA parameters are 161,218,560 (FP32 tensors: 600, base model is the 5-billion parameter scale Wan2.2-TI2V-5B)<\/li>\n<li>Architecture: PEFT LoRA (rank: 64, alpha: 64, dropout: 0.0) targeting 300 Linear layers spanning all 30 transformer blocks (self-attention q\/k\/v\/o, cross-attention q\/k\/v\/o, and FFN 0\/2 of each block)<\/li>\n<li>Output specifications: Resolution of 1280\u00d7704 (704\u00d71280 is also supported per base model specs), with the evaluation profile set to 121 frames at 24 FPS (5-second duration under original model specs)<\/li>\n<li>Recommended steps and sampler: 50 steps, FlowUniPC scheduler (timestep shift: 5.0)<\/li>\n<li>Recommended guidance and weight settings: guidance_scale is 1.0 (CFG=1), and LoRA weight scale is 1.0 (consistent scale during training)<\/li>\n<li>License terms: Distributed under the Apache-2.0 license, allowing use under the specified terms including commercial use<\/li>\n<\/ul>\n<h2>Performance and Quality<\/h2>\n<p>The release source provides the following configuration values regarding the distillation design and inference settings of this adapter:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Setting<\/th>\n<th style=\"text-align: right;\">Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Teacher CFG scale used for distillation<\/td>\n<td style=\"text-align: right;\"><strong>5.0<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Student CFG during training<\/td>\n<td style=\"text-align: right;\"><strong>1.0<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Recommended inference CFG<\/td>\n<td style=\"text-align: right;\"><strong>1.0<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Default LoRA weight scale<\/td>\n<td style=\"text-align: right;\"><strong>1.0<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Teacher \u2192 student denoising steps<\/td>\n<td style=\"text-align: right;\"><strong>50 \u2192 50<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Scheduler \/ timestep shift<\/td>\n<td style=\"text-align: right;\">FlowUniPC \/ <strong>5.0<\/strong><\/td>\n<\/tr>\n<tr>\n<td>LoRA rank \/ alpha<\/td>\n<td style=\"text-align: right;\"><strong>64 \/ 64<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Training checkpoint<\/td>\n<td style=\"text-align: right;\"><strong>step 250<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>In this method, training is performed to minimize the relative guidance-weighted MSE against a fixed teacher model trajectory state, maintaining the native 50 steps identical to the original model. The design does not perform rollout generation by the student model itself during training, and does not include a DMD critic. While it has the advantage of reducing computational load by skipping unconditional branch evaluation, the publisher states that this is a research checkpoint documenting artifact integrity and executed training specifications, and does not guarantee general-purpose quality across all prompts or downstream models.<\/p>\n<p>Regarding the performance of the base Wan2.2-TI2V-5B model itself, it employs the Wan2.2-VAE with a compression rate of 16\u00d716\u00d74 (4\u00d716\u00d716 across temporal, height, and width dimensions), achieving a total compression of 4\u00d732\u00d732 combined with patchify layers. According to the original model&#8217;s published information, it possesses the efficiency to generate 5 seconds of 720P video in under 9 minutes on a single consumer GPU even without optimizations.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>This adapter aims to improve computational efficiency during inference while maintaining the capabilities of the base model Wan2.2-TI2V-5B. Specifically, it streamlines the guidance process via CFG (Classifier-Free Guidance) for both text-to-video and image-to-video tasks.<\/p>\n<p>While conventional CFG-based inference required two forward passes per denoising step\u2014one &#8220;conditional&#8221; and one &#8220;unconditional (negative)&#8221;\u2014applying this LoRA and running with <code>guidance_scale: 1.0<\/code> enables operation with only a single conditional forward pass per step. This is expected to reduce the computational cost of inference while preserving general generation quality trends.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>Original precision<\/td>\n<td>18.6GB<\/td>\n<td>22.4GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. This release is an adapter (LoRA etc.); the table shows what the base model <a href=\"https:\/\/huggingface.co\/Wan-AI\/Wan2.2-TI2V-5B\">Wan-AI\/Wan2.2-TI2V-5B<\/a> needs. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p>The publisher distributes this model as safetensors.<\/p>\n<p><strong>License \u2014 <code>apache-2.0<\/code> (Commercial use allowed):<\/strong> Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats and the license field. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:files --><\/p>\n<h2>Distributed Files<\/h2>\n<p><em>Weight files published in <a href=\"https:\/\/huggingface.co\/Efficient-Large-Model\/LongLive-Plug-Wan2.2-TI2V-5B-cfg\/tree\/main\">Efficient-Large-Model\/LongLive-Plug-Wan2.2-TI2V-5B-cfg<\/a>, listed by this site from the Hugging Face API. Sizes are the actual file sizes.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>File<\/th>\n<th>Size<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>adapter_model.safetensors<\/code><\/td>\n<td>645MB<\/td>\n<\/tr>\n<tr>\n<td><code>generator_lora.pt<\/code><\/td>\n<td>645MB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:files --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model is available from the Hugging Face repository. The following files are provided as distribution formats:<\/p>\n<ul>\n<li><code>adapter_model.safetensors<\/code>: Safe and portable format generator LoRA<\/li>\n<li><code>generator_lora.pt<\/code>: LongLive native payload<\/li>\n<li><code>adapter_config.json<\/code>: PEFT settings and target module information<\/li>\n<li><code>training_config.yaml<\/code>: Training configuration details<\/li>\n<li><code>inference_overrides.yaml<\/code>: Inference settings (50 steps, CFG=1)<\/li>\n<li><code>release_metadata.json<\/code>: Metadata<\/li>\n<li><code>provenance.json<\/code>: Provenance information and checksums<\/li>\n<li><code>SHA256SUMS<\/code>: Checksums for artifacts<\/li>\n<\/ul>\n<p>You can download <code>generator_lora.pt<\/code> using the Hugging Face <code>huggingface_hub<\/code> library with the following command:<\/p>\n<pre><code class=\"language-python\">from huggingface_hub import hf_hub_download\n\nlora_path = hf_hub_download(\n    repo_id=&quot;Efficient-Large-Model\/LongLive-Plug-Wan2.2-TI2V-5B-cfg&quot;,\n    filename=&quot;generator_lora.pt&quot;,\n)\nprint(lora_path)\n<\/code><\/pre>\n<p>When using it as an inference path for LongLive, it is recommended to use settings such as <code>sampling_steps: 50<\/code>, <code>guidance_scale: 1.0<\/code>, and <code>lora_weight_scale: 1.0<\/code>.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/29\/longlive-plug-wan2-2-ti2v-5b-few-step-2\/\">LongLive-Plug-Wan2.2-TI2V-5B-few-step: 24GB+ VRAM, File List<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/longlive-plug-wan21-t2v-14b-few-step-2\/\">LongLive-Plug-Wan2.1-T2V-14B-few-step: 80GB+ VRAM, File List<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/longlive-plug-wan2-1-t2v-14b-cfg\/\">LongLive-Plug-Wan2.1-T2V-14B-cfg Video Generation Model: 80GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/h3-to-ltx-latent-adapter-released\/\">H3-to-LTX-Latent-Adapter: 4GB+ VRAM, File List<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Formats this model is available in<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">Safetensors format guide and models<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-efficient-large-model-en\/\">Efficient Large Model (NVIDIA and MIT): models, licenses and articles<\/a><\/li>\n<li><strong>Other models for the same task<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-video\">Other video generation models<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/Efficient-Large-Model\/LongLive-Plug-Wan2.2-TI2V-5B-cfg\">https:\/\/huggingface.co\/Efficient-Large-Model\/LongLive-Plug-Wan2.2-TI2V-5B-cfg<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>A CFG-distilled LoRA adapter for Wan2.2-TI2V-5B that reduces DiT forward passes to one per step, improving inference efficiency.<\/p>\n","protected":false},"author":1,"featured_media":7371,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[2659,724,2661,111,1547,2643,2645,959,2663],"class_list":["post-7372","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-cfg-en","tag-efficient-large-model-en","tag-longlive-en","tag-lora-en","tag-verified","tag-wan2-2-ti2v-en","tag-wan2-2-ti2v-5b-en","tag--en"],"lang":"en","translations":{"en":7372,"ja":7370},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/7372","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=7372"}],"version-history":[{"count":2,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/7372\/revisions"}],"predecessor-version":[{"id":7494,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/7372\/revisions\/7494"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/7371"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=7372"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=7372"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=7372"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}