Two Quantized Versions of Audio-Visual Video Generation Model FastH

At a Glance
| Item | Value |
|---|---|
| Repository | FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Consumer |
| Family guide | MiniMax-H3 guide (6 articles) |
| Publisher guide | MiniMax: models and licenses |
| Published | 2026-10-06 |
| License | other |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
Two quantized versions of the audio-visual video generation model “FastH3 V2" have been released by the FastVideo team. Both models accept text input and output synchronized video with audio. The base model, FastVideo/FastVideo-FastH3-8-Step-V2, has been released on Hugging Face in two simultaneously published versions: FastVideo-FastH3-8-Step-V2-NVFP4-Consumer quantized in NVFP4 format, and FastVideo-FastH3-8-Step-V2-FP8 quantized in FP8 format. Both support 8-step distilled inference and are distributed with the aim of running on a single local GPU.
Specifications
- Base Model Scale: The base FastVideo-FastH3-8-Step-V2 is reported to have 13.9B parameters.
- Architecture: Both models use FastH3 V2 (50-block configuration) for the main Transformer.
- NVFP4 Version: The MLP linear layers inherit calibrated weights and activation scales from the base NVFP4 quantized version
FastVideo-FastH3-8-Step-V2-NVFP4, while the attention projection layers and sparse attention gates are also kept in NVFP4. – FP8 Version: The attention and MLP linear layers are in FP8 E4M3 format with one scale per output channel, and activations are quantized per-token at runtime. All other parts remain in BF16. - Text Encoder: Both models use an NVFP4-quantized Qwen3-VL restricted to the 50 layers read by H3. In the FP8 version, layers are reportedly dequantized per layer on non-FP4-compatible GPUs.
- VAE: Uses the LynnReal lightweight VAE for video (INT8 weights by Kijai) and H3’s audio VAE for audio.
- Sampling: An 8-step DMD schedule is included in
fastvideo_inference.json, which FastVideo loads automatically. - Target GPUs: The NVFP4 version targets single Blackwell-generation GPUs such as the RTX 5090 and RTX PRO 6000, or DGX Spark. The FP8 version is intended for RTX 4090 and RTX 3090 GPUs without FP4 tensor cores, as well as 16GB and 12GB memory tiers.
- License: The repository listing for both models shows “other", but the base model is said to inherit the MiniMax H3 Community License.
Performance and Quality
According to the base model card, this model (FastH3 V2) is a step-1300 checkpoint trained using data-free DMD2 and VSA-H3 (80% sparse). Regarding the scope of this checkpoint, the publishers themselves state that “text-to-audio-video generation is supported, but FL2VA and Ref2VA are not included in the distillation targets," and further note that “difficult motions, fine details, and certain audio may fall below the quality of the base MiniMax H3 itself." No quantitative benchmark figures are provided in this release.
Strengths and Use Cases
These models excel at text-to-audio-video generation, taking a text prompt as input and simultaneously outputting video with synchronized audio. A major feature is that generation is completed in a small inference process of only 8 steps, making them suitable for local environments where users want to generate audio-visual videos in a short time.
The two newly introduced quantized models are designed to be chosen according to the user’s hardware environment and memory tier:
FastVideo-FastH3-8-Step-V2-NVFP4-Consumer: Suited for text-to-video and audio generation tasks on single Blackwell-generation GPUs like the RTX 5090 or RTX PRO 6000, or environments such as DGX Spark. To operate efficiently on a single GPU, it maintains the 50-block transformer structure while being quantized and packaged in NVFP4 format.FastVideo-FastH3-8-Step-V2-FP8: Intended for use on existing GPU environments without FP4 tensor cores, such as the RTX 4090 or RTX 3090, as well as memory tiers of 16GB and 12GB. Quantization to FP8 precision suppresses memory consumption while enabling local audio-visual video generation.
Note that as a specification of the base model, FL2VA (frame-to-video-and-audio) and Ref2VA (generation from reference images) are not included in the distillation targets, and supported capabilities are limited to text-to-video generation. Additionally, please keep in mind that complex motions, fine detail rendering, and the quality of certain audio may be lower compared to the base MiniMax H3 itself.
How It Differs from Similar Models
“FastH3 Trim", covered in our previous article [[T0]], was an experimental model that aimed to accelerate and reduce size by removing 8 less impactful blocks from the 50 blocks in the FastH3 8-Step V2 transformer to reduce it to 42 layers. In contrast, the models released this time apply NVFP4 and FP8 quantization while maintaining the 50-block structure. They differ in their approach, aiming to run on single GPUs or specific memory tiers without pruning the structure itself.
Additionally, the models covered in [[T1]] were LoRA adapters applied post-hoc to the MiniMax-H3 base model for CFG distillation and acceleration. In contrast, the current models are standalone checkpoints obtained by directly quantizing the distilled model FastVideo-FastH3-8-Step-V2, differing in that no separate adapter needs to be applied.
Distributed Files
Weight files published in FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4-Consumer, listed by this site from the Hugging Face API. Sizes are the actual file sizes.
| File | Size |
|---|---|
audio_vae/diffusion_pytorch_model.safetensors |
605MB |
text_encoder/model-*-of-00004.safetensors |
16.46GB (4 split files) |
transformer/diffusion_pytorch_model-*-of-00006.safetensors |
27.71GB (6 split files) |
transformer/nvfp4_weights.safetensors |
11.92GB |
vae/diffusion_pytorch_model.safetensors |
4.23GB |
vae/minimax_h3_video_vae_int8_convrot.safetensors |
2.14GB |
How to Get It
Both models are available in safetensors format on Hugging Face. They are not Gated models requiring agreement to terms of service for downloading.
Execution uses the FastVideo repository. First, use uv to install the required packages and the VSA-H3 attention kernel:
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
uv venv --python 3.12 --seed
source.venv/bin/activate
UV_TORCH_BACKEND=cu130 uv pip install \
--no-sources-package fastvideo-kernel \
-e ".[fasth3]"
Once installation is complete, run using the included inference script as follows:
python examples/inference/basic/basic_fasth3_8step.py \
--prompt "your prompt" \
--no-warmup \
--repeats 1
During inference, the 8-step DMD sampling schedule (fastvideo_inference.json) is loaded automatically by FastVideo. When running in other multi-GPU environments, following the official guide, adding the --no-replicated-dit --vsa-kernel triton --no-fa4 options is recommended. In that case, the number of GPUs used must be divisible by the number of H3 attention heads (56).
Can You Run It Locally?
The publisher distributes this model as safetensors.
License — other: A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.
Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Our Own Measurements
Measurements We Did Not Take
We have not confirmed that this site meets the commercial-use terms of this model’s license (other). Because this site carries advertising, we did not run the model (no CPU run, answers, quantization comparison, conversion or generation).
Related Articles
- FastH3 Trim: 8-Step Text-to-Video Model with Audio Support
- MiniMax-H3-Character-Swap-LoRA Character Swap LoRA Adapter: File List
- MiniMax-H3 Audio-Visual Video Generation Model: 32GB+ VRAM, File List
What to Read Next
- Explore the same model family → MiniMax-H3 family overview (6 articles, 2 converted builds)
- Formats this model is available in → NVFP4 / MXFP4 format guide and models
- Learn about the publisher → MiniMax: models, licenses and articles
- Other models for the same task → Other video generation models

