LongLive-Plug-Wan2.1-T2V-14B-few-step: 80GB+ VRAM, File List

LongLive-Plug-Wan2.1-T2V-14B-few-step: 80GB+ VRAM, File List

At a Glance

Item Value
Repository Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-few-step
Publisher guide Efficient Large Model (NVIDIA and MIT): models and licenses
Published 2026-09-30
License apache-2.0
Formats safetensors
Source type Primary source (the publisher itself)

Values determined by this site’s code at collection time. Dates are JST.

Overview

Efficient-Large-Model has released a LoRA adapter named LongLive-Plug-Wan2.1-T2V-14B-few-step designed to accelerate video generation, using Wan-AI/Wan2.1-T2V-14B as the base model. This model is designed to enable generation in fewer steps during the text-to-video process. Because this model functions as an adapter, it requires the corresponding base model to operate. Additionally, it is recommended to use it in combination with the corresponding CFG LoRA (LongLive-Plug-Wan2.1-T2V-14B-cfg).

Specifications

This model targets a base model with the following specifications:

  • Base Model: Wan-AI/Wan2.1-T2V-14B
  • Architecture: Video Diffusion DiT (Flow Matching framework)
  • Base Model (14B) Architecture Details:
    • Dimension: 5120 – Input Dimension: 16 – Output Dimension: 16 – Feedforward Dimension: 13824 – Frequency Dimension: 256 – Number of Heads: 40 – Number of Layers: 40
  • Recommended Settings:
    • It is recommended to set the weight ratio of few-step to CFG at 1 : 0.5. Note that this refers to the adapter weight ratio, not the CFG scale during inference.
  • License: apache-2.0

Performance and Quality

The base model Wan2.1-T2V-14B, which serves as the foundation for this adapter, is reported to establish new State-of-the-Art (SOTA) performance benchmarks across both open-source and commercial models.

The detailed configuration of the base model’s architecture is as follows:

→ Scroll horizontally to see all columns

Model Dimension Input Dimension Output Dimension Feedforward Dimension Frequency Dimension Number of Heads Number of Layers
14B 5120 16 16 13824 256 40 40

In evaluations conducted by the publishers, multidimensional verification spanning 14 primary dimensions and 26 sub-dimensions was performed using 1,035 internal prompts. These results were calculated using human preference-based weighting, and it has been confirmed that the model exhibits superior performance compared to existing models.

Additionally, the following features are cited as technologies supporting the quality of the base model:

  • High-Quality Visuals and Motion Dynamics: Capable of generating videos with high-quality visuals and significant motion dynamics.
  • Advanced Text Generation: It is reported to be the first video model capable of generating both Chinese and English text within videos.
  • High-Performance Video VAE: Adopts a proprietary Wan-VAE, making it possible to efficiently encode and decode 1080P videos of arbitrary length while preserving temporal information.
  • Resolution Flexibility: Supports video generation at both 480P and 720P resolutions.

This adapter, LongLive-Plug-Wan2.1-T2V-14B-few-step, achieves efficiency through a reduced number of generation steps by applying distillation techniques to these powerful base models.

Strengths and Use Cases

This model excels at performing fast video generation with a reduced number of steps while maintaining the performance of the base model Wan-AI/Wan2.1-T2V-14B in text-to-video tasks. It is designed as an adapter leveraging PEFT, LoRA, and distillation techniques, and is intended to be used in combination with the corresponding CFG LoRA. Furthermore, it contributes to efficient video generation by taking advantage of the base model’s capabilities, such as generating both Chinese and English text, visual expressions with high motion dynamics, and 480P or 720P resolutions.

How It Differs from Similar Models

While LongLive-Plug-Wan2.1-T2V-14B-cfg, which has already been introduced on this site, aims to perform distillation for conditional-only video generation using Classifier-Free Guidance (CFG) to reduce the execution of unconditional inference branches, this new LongLive-Plug-Wan2.1-T2V-14B-few-step differs in that it functions as a LoRA adapter designed to achieve video generation in fewer steps. Both target the same base model, Wan-AI/Wan2.1-T2V-14B, and it is recommended to use the few-step and CFG adapter weights together at a ratio of 1 : 0.5.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 14.3B parameters (taken from the base model Wan-AI/Wan2.1-T2V-14B)

Your VRAM Quantization File size Est. memory needed
80GB class (A100 / H100) F32 53.2GB 63.9GB

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. This release is an adapter (LoRA etc.); the table shows what the base model Wan-AI/Wan2.1-T2V-14B needs. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

The publisher distributes this model as safetensors.

License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.

Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Distributed Files

Weight files published in Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-few-step, listed by this site from the Hugging Face API. Sizes are the actual file sizes.

File Size
generator_lora_lightx2v.safetensors 1.23GB
training_checkpoint/model.pt 4.91GB

How to Get It

It is available from the official Hugging Face repository. It is distributed as a LoRA adapter weight compatible with the PEFT library, to be used in combination with the corresponding base model and CFG LoRA.

Related Articles

What to Read Next

Sources