LongLive-Plug-Wan2.2-TI2V-5B-cfg Video Generation Model: 24GB+ VRAM

September 30, 2026

LongLive-Plug-Wan2.2-TI2V-5B-cfg Video Generation Model: 24GB+ VRAM

At a Glance

Item Value
Repository Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-cfg
Publisher guide Efficient Large Model (NVIDIA and MIT): models and licenses
Published 2026-09-29
License apache-2.0
Formats safetensors
Source type Primary source (the publisher itself)

Values determined by this site’s code at collection time. Dates are JST.

Overview

LongLive-Plug-Wan2.2-TI2V-5B-cfg has been released as a LoRA adapter applying CFG (Classifier-Free Guidance) distillation for the text- and image-to-video generation model Wan2.2-TI2V-5B.

This model is an adapter that distills the standard CFG=5 guided flow of the base model into a single conditional forward pass at CFG=1 on the student model side. It eliminates the need to calculate unconditional (negative) branches at each denoising step, reducing the DiT (Diffusion Transformer) forward processing to just once per step. Note that denoising schedule distillation is not performed, making this different from DMD (Distribution Matching Distillation) checkpoints that carry out few-step generation.

Specifications

The specifications and recommended settings listed in the model card are as follows:

  • Parameter count: LoRA parameters are 161,218,560 (FP32 tensors: 600, base model is the 5-billion parameter scale Wan2.2-TI2V-5B)
  • Architecture: PEFT LoRA (rank: 64, alpha: 64, dropout: 0.0) targeting 300 Linear layers spanning all 30 transformer blocks (self-attention q/k/v/o, cross-attention q/k/v/o, and FFN 0/2 of each block)
  • Output specifications: Resolution of 1280×704 (704×1280 is also supported per base model specs), with the evaluation profile set to 121 frames at 24 FPS (5-second duration under original model specs)
  • Recommended steps and sampler: 50 steps, FlowUniPC scheduler (timestep shift: 5.0)
  • Recommended guidance and weight settings: guidance_scale is 1.0 (CFG=1), and LoRA weight scale is 1.0 (consistent scale during training)
  • License terms: Distributed under the Apache-2.0 license, allowing use under the specified terms including commercial use

Performance and Quality

The release source provides the following configuration values regarding the distillation design and inference settings of this adapter:

Setting Value
Teacher CFG scale used for distillation 5.0
Student CFG during training 1.0
Recommended inference CFG 1.0
Default LoRA weight scale 1.0
Teacher → student denoising steps 50 → 50
Scheduler / timestep shift FlowUniPC / 5.0
LoRA rank / alpha 64 / 64
Training checkpoint step 250

In this method, training is performed to minimize the relative guidance-weighted MSE against a fixed teacher model trajectory state, maintaining the native 50 steps identical to the original model. The design does not perform rollout generation by the student model itself during training, and does not include a DMD critic. While it has the advantage of reducing computational load by skipping unconditional branch evaluation, the publisher states that this is a research checkpoint documenting artifact integrity and executed training specifications, and does not guarantee general-purpose quality across all prompts or downstream models.

Regarding the performance of the base Wan2.2-TI2V-5B model itself, it employs the Wan2.2-VAE with a compression rate of 16×16×4 (4×16×16 across temporal, height, and width dimensions), achieving a total compression of 4×32×32 combined with patchify layers. According to the original model’s published information, it possesses the efficiency to generate 5 seconds of 720P video in under 9 minutes on a single consumer GPU even without optimizations.

Strengths and Use Cases

This adapter aims to improve computational efficiency during inference while maintaining the capabilities of the base model Wan2.2-TI2V-5B. Specifically, it streamlines the guidance process via CFG (Classifier-Free Guidance) for both text-to-video and image-to-video tasks.

While conventional CFG-based inference required two forward passes per denoising step—one “conditional" and one “unconditional (negative)"—applying this LoRA and running with guidance_scale: 1.0 enables operation with only a single conditional forward pass per step. This is expected to reduce the computational cost of inference while preserving general generation quality trends.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch)

Your VRAM Quantization File size Est. memory needed
24GB (RTX 4090 / 3090, etc.) Original precision 18.6GB 22.4GB

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. This release is an adapter (LoRA etc.); the table shows what the base model Wan-AI/Wan2.2-TI2V-5B needs. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

The publisher distributes this model as safetensors.

License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.

Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Distributed Files

Weight files published in Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-cfg, listed by this site from the Hugging Face API. Sizes are the actual file sizes.

File Size
adapter_model.safetensors 645MB
generator_lora.pt 645MB

How to Get It

This model is available from the Hugging Face repository. The following files are provided as distribution formats:

  • adapter_model.safetensors: Safe and portable format generator LoRA
  • generator_lora.pt: LongLive native payload
  • adapter_config.json: PEFT settings and target module information
  • training_config.yaml: Training configuration details
  • inference_overrides.yaml: Inference settings (50 steps, CFG=1)
  • release_metadata.json: Metadata
  • provenance.json: Provenance information and checksums
  • SHA256SUMS: Checksums for artifacts

You can download generator_lora.pt using the Hugging Face huggingface_hub library with the following command:

from huggingface_hub import hf_hub_download

lora_path = hf_hub_download(
    repo_id="Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-cfg",
    filename="generator_lora.pt",
)
print(lora_path)

When using it as an inference path for LongLive, it is recommended to use settings such as sampling_steps: 50, guidance_scale: 1.0, and lora_weight_scale: 1.0.

Related Articles

What to Read Next

Sources