LongLive-Plug-Wan2.2-TI2V-5B-few-step: 24GB+ VRAM, File List

At a Glance
| Item | Value |
|---|---|
| Repository | Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-few-step |
| Publisher guide | Efficient Large Model (NVIDIA and MIT): models and licenses |
| Published | 2026-09-29 |
| License | apache-2.0 |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code at collection time. Dates are JST.
Overview
Efficient-Large-Model has released a Generator LoRA, LongLive-Plug-Wan2.2-TI2V-5B-few-step, designed for video generation models based on Wan-AI/Wan2.2-TI2V-5B to enable inference in an extremely low number of steps. This LoRA supports both Text-to-Video and Image-to-Video tasks, which output videos from text or image inputs, achieving a 4-step non-autoregressive full-sequence inference.
The base model, Wan2.2-TI2V-5B, supports 720P resolution at 24fps video generation and is among the fastest available 720P@24fps models. The checkpoint in this release is a trained model at iteration 1500 from dev/nonAR experiments using 16 GPUs. real_guidance_scale=3 was set during training. Note that this distilled Generator has been validated with guidance_scale=1.0, which does not use CFG (Classifier-Free Guidance) during inference. The published file is a lossless extraction of the Generator only. While the original checkpoint also includes a training-only Critic LoRA, only the Generator LoRA is provided in this release.
Specifications
- Base Model:
Wan-AI/Wan2.2-TI2V-5B - Architecture: full-sequence non-AR
- Training Objective: DMD (Differentiable Diffusion Model) with backward simulation
- Real Guidance Scale (RGS) during Training: 3
- Checkpoint Identifier:
checkpoint_model_001500 - World Size during Training: 16
- LoRA Rank / Alpha: 128 / 128
- Generator LoRA Parameters: 322,437,120 (600 FP32 tensors)
- Denoising Steps during Inference: 4 steps
- CFG Scale during Inference: 1.0 (No CFG)
- License: Apache-2.0
Strengths and Use Cases
This model is specialized in dramatically improving generation speed while maintaining the advanced video generation capabilities of the base model, Wan2.2-TI2V-5B. Its main use cases are both Text-to-Video and Image-to-Video generation. The primary advantage is that it can complete high-quality video generation, which normally requires many steps, in just 4 steps of inference.
The Wan2.2 series, which serves as the base model, features a significantly expanded training dataset compared to the previous version, Wan2.1. Specifically, image data has increased by 65.6% and video data by 83.2%, which greatly improves versatility in depicting complex motions, semantics, and visual aesthetics. It is developed with a particular focus on “cinematic aesthetics" and is trained on datasets with detailed labels for lighting, composition, contrast, and color tone. Therefore, it excels at generating precise and controllable cinematic-style videos tailored to user preferences.
Additionally, the Wan2.2-VAE adopted by this model achieves a high compression rate of 16×16×4, enabling efficient generation while maintaining high-quality video reconstruction. This allows for the generation of 720P resolution (1280×704 or 704×1280) videos at 24fps. Because this model natively supports both text and image inputs within a single framework, advanced control combining prompt instructions and image-based composition specification is possible.
Applying this LoRA makes it a powerful tool for creative applications requiring many trial-and-error iterations in a short time, applications demanding near-real-time responses, or technical environments aiming to obtain high-quality videos while suppressing computational resources. By adopting non-autoregressive (non-AR) full-sequence inference, the efficiency of the entire generation process is optimized.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch)
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 24GB (RTX 4090 / 3090, etc.) | Original precision | 18.6GB | 22.4GB |
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. This release is an adapter (LoRA etc.); the table shows what the base model Wan-AI/Wan2.2-TI2V-5B needs. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
The publisher distributes this model as safetensors.
License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.
Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Distributed Files
Weight files published in Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-few-step, listed by this site from the Hugging Face API. Sizes are the actual file sizes.
| File | Size |
|---|---|
adapter_model.safetensors |
1.29GB |
generator_lora.pt |
1.29GB |
How to Get It
LoRA weights and related configuration files for this model are available from the Hugging Face repository Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-few-step. This repository holds selected training checkpoint weights, configuration files, runtime source code, and license notices.
The main distributed files are as follows:
– generator_lora.pt: Native format weights for use in the LongLive wrapper
– adapter_model.safetensors: Safetensors format weights containing 600 tensors
– adapter_config.json: Configuration file including PEFT Rank/Alpha settings
– inference_config.yaml: Inference configuration file including 4-step inference and No-CFG settings
– training_config.yaml: Detailed configuration records during training with 16 GPUs
– provenance.json: Checksum records of source files and published files
When downloading weights using Python, the following code can be used:
from huggingface_hub import hf_hub_download
lora_path = hf_hub_download(
repo_id="Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-few-step",
filename="generator_lora.pt",
)
During inference, the following settings must be strictly observed. These are settings dedicated to this adapter, and default settings for general autoregressive models should not be used:
– Sampling Steps: 4
– Guidance Scale during Inference: 1.0 (No CFG)
– LoRA Rank / Alpha: 128 / 128
– generator_is_causal: false (non-autoregressive setting)
This model is released under the Apache-2.0 license, allowing flexible use for both commercial and research purposes. No special consent process (Gated) is required to obtain it.
Related Articles
- LongLive-Plug-Wan2.2-TI2V-5B-cfg Video Generation Model: 24GB+ VRAM
- LongLive-Plug-Wan2.1-T2V-14B-few-step: 80GB+ VRAM, File List
- LongLive-Plug-Wan2.1-T2V-14B-cfg Video Generation Model: 80GB+ VRAM
- H3-to-LTX-Latent-Adapter: 4GB+ VRAM, File List
What to Read Next
- Formats this model is available in → Safetensors format guide and models
- Learn about the publisher → Efficient Large Model (NVIDIA and MIT): models, licenses and articles
- Other models for the same task → Other video generation models

