LongLive-Plug-Wan2.2-TI2V-5B-cfg Video Generation Model: 24GB+ VRAM

At a Glance
| Item | Value |
|---|---|
| Repository | Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-cfg |
| Publisher guide | Efficient Large Model (NVIDIA and MIT): models and licenses |
| Published | 2026-09-29 |
| License | apache-2.0 |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code at collection time. Dates are JST.
Overview
LongLive-Plug-Wan2.2-TI2V-5B-cfg has been released as a LoRA adapter applying CFG (Classifier-Free Guidance) distillation for the text- and image-to-video generation model Wan2.2-TI2V-5B.
This model is an adapter that distills the standard CFG=5 guided flow of the base model into a single conditional forward pass at CFG=1 on the student model side. It eliminates the need to calculate unconditional (negative) branches at each denoising step, reducing the DiT (Diffusion Transformer) forward processing to just once per step. Note that denoising schedule distillation is not performed, making this different from DMD (Distribution Matching Distillation) checkpoints that carry out few-step generation.
Specifications
The specifications and recommended settings listed in the model card are as follows:
- Parameter count: LoRA parameters are 161,218,560 (FP32 tensors: 600, base model is the 5-billion parameter scale Wan2.2-TI2V-5B)
- Architecture: PEFT LoRA (rank: 64, alpha: 64, dropout: 0.0) targeting 300 Linear layers spanning all 30 transformer blocks (self-attention q/k/v/o, cross-attention q/k/v/o, and FFN 0/2 of each block)
- Output specifications: Resolution of 1280×704 (704×1280 is also supported per base model specs), with the evaluation profile set to 121 frames at 24 FPS (5-second duration under original model specs)
- Recommended steps and sampler: 50 steps, FlowUniPC scheduler (timestep shift: 5.0)
- Recommended guidance and weight settings: guidance_scale is 1.0 (CFG=1), and LoRA weight scale is 1.0 (consistent scale during training)
- License terms: Distributed under the Apache-2.0 license, allowing use under the specified terms including commercial use
Performance and Quality
The release source provides the following configuration values regarding the distillation design and inference settings of this adapter:
| Setting | Value |
|---|---|
| Teacher CFG scale used for distillation | 5.0 |
| Student CFG during training | 1.0 |
| Recommended inference CFG | 1.0 |
| Default LoRA weight scale | 1.0 |
| Teacher → student denoising steps | 50 → 50 |
| Scheduler / timestep shift | FlowUniPC / 5.0 |
| LoRA rank / alpha | 64 / 64 |
| Training checkpoint | step 250 |
In this method, training is performed to minimize the relative guidance-weighted MSE against a fixed teacher model trajectory state, maintaining the native 50 steps identical to the original model. The design does not perform rollout generation by the student model itself during training, and does not include a DMD critic. While it has the advantage of reducing computational load by skipping unconditional branch evaluation, the publisher states that this is a research checkpoint documenting artifact integrity and executed training specifications, and does not guarantee general-purpose quality across all prompts or downstream models.
Regarding the performance of the base Wan2.2-TI2V-5B model itself, it employs the Wan2.2-VAE with a compression rate of 16×16×4 (4×16×16 across temporal, height, and width dimensions), achieving a total compression of 4×32×32 combined with patchify layers. According to the original model’s published information, it possesses the efficiency to generate 5 seconds of 720P video in under 9 minutes on a single consumer GPU even without optimizations.
Strengths and Use Cases
This adapter aims to improve computational efficiency during inference while maintaining the capabilities of the base model Wan2.2-TI2V-5B. Specifically, it streamlines the guidance process via CFG (Classifier-Free Guidance) for both text-to-video and image-to-video tasks.
While conventional CFG-based inference required two forward passes per denoising step—one “conditional" and one “unconditional (negative)"—applying this LoRA and running with guidance_scale: 1.0 enables operation with only a single conditional forward pass per step. This is expected to reduce the computational cost of inference while preserving general generation quality trends.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch)
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 24GB (RTX 4090 / 3090, etc.) | Original precision | 18.6GB | 22.4GB |
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. This release is an adapter (LoRA etc.); the table shows what the base model Wan-AI/Wan2.2-TI2V-5B needs. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
The publisher distributes this model as safetensors.
License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.
Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Distributed Files
Weight files published in Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-cfg, listed by this site from the Hugging Face API. Sizes are the actual file sizes.
| File | Size |
|---|---|
adapter_model.safetensors |
645MB |
generator_lora.pt |
645MB |
How to Get It
This model is available from the Hugging Face repository. The following files are provided as distribution formats:
adapter_model.safetensors: Safe and portable format generator LoRAgenerator_lora.pt: LongLive native payloadadapter_config.json: PEFT settings and target module informationtraining_config.yaml: Training configuration detailsinference_overrides.yaml: Inference settings (50 steps, CFG=1)release_metadata.json: Metadataprovenance.json: Provenance information and checksumsSHA256SUMS: Checksums for artifacts
You can download generator_lora.pt using the Hugging Face huggingface_hub library with the following command:
from huggingface_hub import hf_hub_download
lora_path = hf_hub_download(
repo_id="Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-cfg",
filename="generator_lora.pt",
)
print(lora_path)
When using it as an inference path for LongLive, it is recommended to use settings such as sampling_steps: 50, guidance_scale: 1.0, and lora_weight_scale: 1.0.
Related Articles
- LongLive-Plug-Wan2.2-TI2V-5B-few-step: 24GB+ VRAM, File List
- LongLive-Plug-Wan2.1-T2V-14B-cfg Video Generation Model: 80GB+ VRAM
- H3-to-LTX-Latent-Adapter: 4GB+ VRAM, File List
What to Read Next
- Formats this model is available in → Safetensors format guide and models
- Learn about the publisher → Efficient Large Model (NVIDIA and MIT): models, licenses and articles
- Other models for the same task → Other video generation models

