LongLive-Plug-Wan2.1-T2V-14B-few-step: 80GB+ VRAM, File List

At a Glance
| Item | Value |
|---|---|
| Repository | Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-few-step |
| Publisher guide | Efficient Large Model (NVIDIA and MIT): models and licenses |
| Published | 2026-09-30 |
| License | apache-2.0 |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code at collection time. Dates are JST.
Overview
Efficient-Large-Model has released a LoRA adapter named LongLive-Plug-Wan2.1-T2V-14B-few-step designed to accelerate video generation, using Wan-AI/Wan2.1-T2V-14B as the base model. This model is designed to enable generation in fewer steps during the text-to-video process. Because this model functions as an adapter, it requires the corresponding base model to operate. Additionally, it is recommended to use it in combination with the corresponding CFG LoRA (LongLive-Plug-Wan2.1-T2V-14B-cfg).
Specifications
This model targets a base model with the following specifications:
- Base Model:
Wan-AI/Wan2.1-T2V-14B - Architecture: Video Diffusion DiT (Flow Matching framework)
- Base Model (14B) Architecture Details:
- Dimension: 5120 – Input Dimension: 16 – Output Dimension: 16 – Feedforward Dimension: 13824 – Frequency Dimension: 256 – Number of Heads: 40 – Number of Layers: 40
- Recommended Settings:
- It is recommended to set the weight ratio of
few-steptoCFGat1 : 0.5. Note that this refers to the adapter weight ratio, not the CFG scale during inference.
- It is recommended to set the weight ratio of
- License:
apache-2.0
Performance and Quality
The base model Wan2.1-T2V-14B, which serves as the foundation for this adapter, is reported to establish new State-of-the-Art (SOTA) performance benchmarks across both open-source and commercial models.
The detailed configuration of the base model’s architecture is as follows:
→ Scroll horizontally to see all columns
| Model | Dimension | Input Dimension | Output Dimension | Feedforward Dimension | Frequency Dimension | Number of Heads | Number of Layers |
|---|---|---|---|---|---|---|---|
| 14B | 5120 | 16 | 16 | 13824 | 256 | 40 | 40 |
In evaluations conducted by the publishers, multidimensional verification spanning 14 primary dimensions and 26 sub-dimensions was performed using 1,035 internal prompts. These results were calculated using human preference-based weighting, and it has been confirmed that the model exhibits superior performance compared to existing models.
Additionally, the following features are cited as technologies supporting the quality of the base model:
- High-Quality Visuals and Motion Dynamics: Capable of generating videos with high-quality visuals and significant motion dynamics.
- Advanced Text Generation: It is reported to be the first video model capable of generating both Chinese and English text within videos.
- High-Performance Video VAE: Adopts a proprietary
Wan-VAE, making it possible to efficiently encode and decode 1080P videos of arbitrary length while preserving temporal information. - Resolution Flexibility: Supports video generation at both 480P and 720P resolutions.
This adapter, LongLive-Plug-Wan2.1-T2V-14B-few-step, achieves efficiency through a reduced number of generation steps by applying distillation techniques to these powerful base models.
Strengths and Use Cases
This model excels at performing fast video generation with a reduced number of steps while maintaining the performance of the base model Wan-AI/Wan2.1-T2V-14B in text-to-video tasks. It is designed as an adapter leveraging PEFT, LoRA, and distillation techniques, and is intended to be used in combination with the corresponding CFG LoRA. Furthermore, it contributes to efficient video generation by taking advantage of the base model’s capabilities, such as generating both Chinese and English text, visual expressions with high motion dynamics, and 480P or 720P resolutions.
How It Differs from Similar Models
While LongLive-Plug-Wan2.1-T2V-14B-cfg, which has already been introduced on this site, aims to perform distillation for conditional-only video generation using Classifier-Free Guidance (CFG) to reduce the execution of unconditional inference branches, this new LongLive-Plug-Wan2.1-T2V-14B-few-step differs in that it functions as a LoRA adapter designed to achieve video generation in fewer steps. Both target the same base model, Wan-AI/Wan2.1-T2V-14B, and it is recommended to use the few-step and CFG adapter weights together at a ratio of 1 : 0.5.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 14.3B parameters (taken from the base model Wan-AI/Wan2.1-T2V-14B)
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 80GB class (A100 / H100) | F32 | 53.2GB | 63.9GB |
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. This release is an adapter (LoRA etc.); the table shows what the base model Wan-AI/Wan2.1-T2V-14B needs. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
The publisher distributes this model as safetensors.
License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.
Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Distributed Files
Weight files published in Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-few-step, listed by this site from the Hugging Face API. Sizes are the actual file sizes.
| File | Size |
|---|---|
generator_lora_lightx2v.safetensors |
1.23GB |
training_checkpoint/model.pt |
4.91GB |
How to Get It
It is available from the official Hugging Face repository. It is distributed as a LoRA adapter weight compatible with the PEFT library, to be used in combination with the corresponding base model and CFG LoRA.
- Repository: Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-few-step
- Distribution Format:
safetensors
Related Articles
- LongLive-Plug-Wan2.1-T2V-14B-cfg Video Generation Model: 80GB+ VRAM
- LongLive-Plug-Wan2.2-TI2V-5B-cfg Video Generation Model: 24GB+ VRAM
- LongLive-Plug-Wan2.2-TI2V-5B-few-step: 24GB+ VRAM, File List
- H3-to-LTX-Latent-Adapter: 4GB+ VRAM, File List
What to Read Next
- Formats this model is available in → Safetensors format guide and models
- Learn about the publisher → Efficient Large Model (NVIDIA and MIT): models, licenses and articles
- Other models for the same task → Other video generation models
