{"id":10202,"date":"2026-10-06T04:09:37","date_gmt":"2026-10-05T19:09:37","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/06\/fasth3-trim-8-step-video-generation-model\/"},"modified":"2026-10-06T12:09:14","modified_gmt":"2026-10-06T03:09:14","slug":"fasth3-trim-8-step-video-generation-model","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/06\/fasth3-trim-8-step-video-generation-model\/","title":{"rendered":"FastH3 Trim: 8-Step Text-to-Video Model with Audio Support"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/FastVideo\/FastVideo-FastH3-Trim-8-Step-NVFP4\">FastVideo\/FastVideo-FastH3-Trim-8-Step-NVFP4<\/a><\/td>\n<\/tr>\n<tr>\n<td>Family guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-minimaxai-minimax-h3-en\/\">MiniMax-H3 guide (6 articles)<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-minimax-en\/\">MiniMax: models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-03<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>other<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>safetensors<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>From the FastVideo team comes &#8220;FastH3 Trim&#8221;, an 8-step video generation model capable of producing text-to-video content with synchronized audio. Alongside the BF16 edition <code>FastVideo\/FastVideo-FastH3-Trim-8-Step<\/code>, quantized variants <code>FastVideo\/FastVideo-FastH3-Trim-8-Step-NVFP4<\/code> and <code>FastVideo\/FastVideo-FastH3-Trim-8-Step-FP8<\/code> have been released simultaneously.<\/p>\n<p>FastH3 Trim is an experimental lightweight model based on <code>FastVideo\/FastVideo-FastH3-8-Step-V2<\/code>, featuring reduced transformer blocks and various optimizations. By taking text input, it generates synchronized audio and video in a mere 8-step inference process. While block reduction aims to increase speed and reduce size, it is explicitly noted that there is a quality trade-off involved.<\/p>\n<h2>Specifications<\/h2>\n<p>Specifications confirmed from the released model cards are as follows:<\/p>\n<ul>\n<li><strong>Architecture and Structure:<\/strong><\/li>\n<li>Out of the 50 transformer blocks in MiniMax H3, the 8 blocks with the smallest impact on video and audio prediction results were removed, resulting in a total of 42 blocks. &#8211; The AdaLN timestep projection in each block has been replaced with a shared rank-16 basis. &#8211; Trained using 8-step DMD2 (checkpoint 300).<\/li>\n<li><strong>Sampling and Inference Specifications:<\/strong><\/li>\n<li>DMD steps: 8 steps (timesteps: 999, 874, 749, 624, 500, 375, 250, 125). &#8211; Sparse attention: Configured to maintain 20% of tiles (80% sparsity). &#8211; The inference schedule is kept in <code>fastvideo_inference.json<\/code> and is automatically loaded by the &#8220;FastVideo&#8221; inference framework.<\/li>\n<li><strong>Component Modules:<\/strong><\/li>\n<li>Text encoder: Adopts the NVFP4 format Qwen3-VL trimmed to 50 layers as referenced by H3. &#8211; VAE: Uses the lightweight video VAE &#8220;LynnReal&#8221; with INT8 weights applied by Kijai, along with the H3 audio VAE. &#8211; Diffusers pipeline: Compatible with <code>MiniMaxH3ModularPipeline<\/code>.<\/li>\n<li><strong>Variant Specifications:<\/strong><\/li>\n<li><strong>BF16 version (<code>FastVideo-FastH3-Trim-8-Step<\/code>):<\/strong> Transformer size is 34.9 GiB (AdaLN coefficients are in FP16). Intended as a source for quantization releases or for local MLX conversion on Apple Silicon. &#8211; <strong>NVFP4 version (<code>FastVideo-FastH3-Trim-8-Step-NVFP4<\/code>):<\/strong> Transformer size is 11.1 GiB. Attention, MLP, and sparse attention gated linear layers are composed of NVFP4, featuring static activation scales calibrated with 1,000 prompts across all 8 steps. Targeted for RTX 5090, RTX PRO 6000, and DGX Spark. &#8211; <strong>FP8 version (<code>FastVideo-FastH3-Trim-8-Step-FP8<\/code>):<\/strong> Transformer size is 19.9 GiB. Attention and MLP linear layers consist of FP8 E4M3 with 1 scale per output channel, with activations quantized per-token at runtime. Designed to run on the RTX 4090, with layerwise offload support enabling operation on 16 GB and 12 GB environments.<\/li>\n<li><strong>License:<\/strong><\/li>\n<li>Designated as <code>other<\/code> (the original model description states it inherits the MiniMax H3 Community License).<\/li>\n<\/ul>\n<h2>Performance and Quality<\/h2>\n<p>While benchmark score tables are not included in the model card, the publishers have detailed several characteristics regarding quality and generation behavior.<\/p>\n<p>First, as a characteristic of this model itself, it is reported that while the removal of transformer blocks achieves a smaller and faster model, it comes with a certain drop in quality. Consequently, the publishers recommend using the full-quality original model <code>FastVideo-FastH3-8-Step-V2<\/code> for use cases where generation quality is the top priority.<\/p>\n<p>Additionally, regarding the performance traits of the parent model FastH3-8-Step-V2, it is distilled specifically for text-to-audio-video generation, meaning generation from reference images or videos (FL2VA and Ref2VA) is not distilled. The publishers note that compared to the original base MiniMax H3, quality may be lower in complex motions, fine details, and certain audio expressions.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>FastH3 Trim specializes in running synchronized text-to-video generation with audio in local environments constrained by hardware resources or requiring high speeds.<\/p>\n<p>Its greatest strength is the high inference efficiency achieved by combining transformer block pruning (removing 8 out of 50 blocks), 8-step DMD2 distillation, and sparse attention. Compared to traditional video generation models requiring multi-step inference, the number of steps is significantly reduced, making it suitable for rapid preview generation, prototyping, and experimental video and audio generation on local PCs.<\/p>\n<p>Another major feature is the availability of three variants tailored to user execution environments:<\/p>\n<ul>\n<li><strong>BF16 version:<\/strong> Suitable as a source model for developers wanting to perform local MLX conversion on Apple Silicon environments or apply their own quantization processing.<\/li>\n<li><strong>NVFP4 version:<\/strong> Optimized for environments equipped with NVIDIA Blackwell architectures (such as RTX 5090, RTX PRO 6000, and DGX Spark), aimed at high-speed inference leveraging cutting-edge hardware performance.<\/li>\n<li><strong>FP8 version:<\/strong> Suited for execution on existing consumer GPUs like the RTX 4090, supporting local generation attempts even on VRAM 16 GB or 12 GB class GPU environments when combined with layerwise offload functions.<\/li>\n<\/ul>\n<p>As noted by the publishers, due to quality degradation accompanying block pruning, the full-quality version <code>FastH3 V2<\/code> is recommended for scenarios prioritizing absolute visual and audio quality, such as commercial production. Furthermore, since the distillation scope of the parent model is limited to text inputs, it is not suited for use cases requiring generation based on existing videos or reference images (FL2VA and Ref2VA).<\/p>\n<h2>How It Differs from Similar Models<\/h2>\n<p>Specific differences from MiniMax-H3 related models previously covered on our site are as follows:<\/p>\n<p>In our previous coverage of <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/30\/minimax-h3-lora-adapters\/\">Two LoRA Adapters Released for the &#8220;MiniMax-H3&#8221; Video Generation Model<\/a>, LoRA adapters were provided to add CFG (classifier-free guidance) distillation and fewer steps to the base MiniMax-H3 model. In contrast, FastH3 Trim is not an adapter format, but a standalone lightweight model that underwent distillation training after physically pruning the transformer blocks of the model itself. It differs by achieving a compact model size and fast 8-step inference on its own, without requiring additional adapters applied to a base model.<\/p>\n<p>Additionally, the model covered in <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/akatz-ai-minimax-h3-character-swap-lora\/\">&#8220;MiniMax-H3-Character-Swap-LoRA&#8221; Video Generation Model: List of Distributed Files<\/a> was a specialized adapter dedicated to the Video-to-Video (Ref2VA) task of taking existing video inputs and swapping characters. Meanwhile, FastH3 Trim specializes in Text-to-Video, generating synchronized audio and video from scratch, and functions such as Ref2VA are excluded from distillation. The purpose and approach are clearly distinct, focusing on reducing execution load and boosting inference speed for the basic text-to-generation process itself rather than applying task-specific adapters.<\/p>\n<p><!-- lmw:files --><\/p>\n<h2>Distributed Files<\/h2>\n<p><em>Weight files published in <a href=\"https:\/\/huggingface.co\/FastVideo\/FastVideo-FastH3-Trim-8-Step\/tree\/main\">FastVideo\/FastVideo-FastH3-Trim-8-Step<\/a>, listed by this site from the Hugging Face API. Sizes are the actual file sizes.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>File<\/th>\n<th>Size<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>audio_vae\/diffusion_pytorch_model.safetensors<\/code><\/td>\n<td>605MB<\/td>\n<\/tr>\n<tr>\n<td><code>text_encoder\/model-*-of-00004.safetensors<\/code><\/td>\n<td>16.46GB (4 split files)<\/td>\n<\/tr>\n<tr>\n<td><code>transformer\/diffusion_pytorch_model-*-of-00008.safetensors<\/code><\/td>\n<td>37.52GB (8 split files)<\/td>\n<\/tr>\n<tr>\n<td><code>vae\/diffusion_pytorch_model.safetensors<\/code><\/td>\n<td>4.23GB<\/td>\n<\/tr>\n<tr>\n<td><code>vae\/minimax_h3_video_vae_int8_convrot.safetensors<\/code><\/td>\n<td>2.14GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:files --><\/p>\n<h2>How to Get It<\/h2>\n<p>Weights for each model are published on Hugging Face in safetensors format. Since it is not a gated model, you can download them directly without prior application or consent procedures.<\/p>\n<p>Depending on your use case, they can be obtained from the following repositories:<\/p>\n<ul>\n<li>BF16 source weights: <code>FastVideo\/FastVideo-FastH3-Trim-8-Step<\/code><\/li>\n<li>NVFP4 quantized version: <code>FastVideo\/FastVideo-FastH3-Trim-8-Step-NVFP4<\/code><\/li>\n<li>FP8 quantized version: <code>FastVideo\/FastVideo-FastH3-Trim-8-Step-FP8<\/code><\/li>\n<\/ul>\n<p>In addition to supporting Diffusers (pipeline: <code>MiniMaxH3ModularPipeline<\/code>) as a framework, you can perform inference using the FastVideo development repository. The inference schedule is saved in <code>fastvideo_inference.json<\/code> within each repository and is automatically loaded by FastVideo.<\/p>\n<p>The setup procedure using the FastVideo repository is outlined as follows:<\/p>\n<pre><code class=\"language-bash\">git clone https:\/\/github.com\/hao-ai-lab\/FastVideo.git\ncd FastVideo\nuv venv --python 3.12 --seed\nsource.venv\/bin\/activate\nUV_TORCH_BACKEND=cu130 uv pip install \\\n  --no-sources-package fastvideo-kernel \\\n  -e &quot;.[fasth3]&quot;\n<\/code><\/pre>\n<p>To run inference, use the included script:<\/p>\n<pre><code class=\"language-bash\">python examples\/inference\/basic\/basic_fasth3_8step.py \\\n  --prompt &quot;your prompt&quot; \\\n  --no-warmup \\\n  --repeats 1\n<\/code><\/pre>\n<p>Note that running the FP8 version in 16 GB or 12 GB environments requires configuring layerwise offload. Furthermore, it is guided that running in multi-GPU environments requires a configuration where the number of GPUs evenly divides 56, which is the number of attention heads in MiniMax H3.<\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p>The publisher distributes this model as safetensors.<\/p>\n<p><strong>License \u2014 <code>other<\/code>:<\/strong> A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats and the license field. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<h3>Measurements We Did Not Take<\/h3>\n<p>We have not confirmed that this site meets the commercial-use terms of this model&#8217;s license (<code>other<\/code>). Because this site carries advertising, we did not run the model (no CPU run, answers, quantization comparison, conversion or generation).<\/p>\n<p><!-- \/lmw:lab --><\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/06\/fasth3-v2-quantized-models-released\/\">Two Quantized Versions of Audio-Visual Video Generation Model FastH<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/minimax-h3-character-swap-lora\/\">MiniMax-H3-Character-Swap-LoRA Character Swap LoRA Adapter: File List<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/minimax-h3-open-omnimodal-video-generation\/\">MiniMax-H3 Audio-Visual Video Generation Model: 32GB+ VRAM, File List<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Explore the same model family<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-minimaxai-minimax-h3-en\/\">MiniMax-H3 family overview (6 articles, 2 converted builds)<\/a><\/li>\n<li><strong>Formats this model is available in<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-fp4-en\/\">NVFP4 \/ MXFP4 format guide and models<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-minimax-en\/\">MiniMax: models, licenses and articles<\/a><\/li>\n<li><strong>Other models for the same task<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-video\">Other video generation models<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/FastVideo\/FastVideo-FastH3-Trim-8-Step\">FastVideo\/FastVideo-FastH3-Trim-8-Step<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/FastVideo\/FastVideo-FastH3-Trim-8-Step-NVFP4\">FastVideo\/FastVideo-FastH3-Trim-8-Step-NVFP4<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/FastVideo\/FastVideo-FastH3-Trim-8-Step-FP8\">FastVideo\/FastVideo-FastH3-Trim-8-Step-FP8<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/FastVideo\/FastVideo-FastH3-8-Step-V2\">FastVideo\/FastVideo-FastH3-8-Step-V2<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/hao-ai-lab\/FastVideo\">FastVideo GitHub Repository<\/a><\/li>\n<li><a href=\"https:\/\/haoailab.com\/blogs\/fasth3-rtx\/\">FastH3 on Consumer Hardware (Blog)<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Explore FastH3 Trim, an 8-step text-to-video model with audio support released by the FastVideo team, available in BF16, NVFP4, and FP8 variants.<\/p>\n","protected":false},"author":1,"featured_media":10201,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[3029,3031,1133,730,1547,959,2663,213],"class_list":["post-10202","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-fasth3-trim-en","tag-fastvideo-en","tag-fp8-en","tag-minimax-h3-en","tag-verified","tag--en"],"lang":"en","translations":{"en":10202,"ja":10200},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10202","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=10202"}],"version-history":[{"count":2,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10202\/revisions"}],"predecessor-version":[{"id":10327,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10202\/revisions\/10327"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/10201"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=10202"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=10202"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=10202"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}