{"id":6107,"date":"2026-09-28T11:09:40","date_gmt":"2026-09-28T02:09:40","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/minimax-h3-character-swap-lora\/"},"modified":"2026-09-28T11:09:40","modified_gmt":"2026-09-28T02:09:40","slug":"minimax-h3-character-swap-lora","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/minimax-h3-character-swap-lora\/","title":{"rendered":"MiniMax-H3-Character-Swap-LoRA Character Swap LoRA Adapter: File List"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/akatz-ai\/MiniMax-H3-Character-Swap-LoRA\">akatz-ai\/MiniMax-H3-Character-Swap-LoRA<\/a><\/td>\n<\/tr>\n<tr>\n<td>Family guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-minimaxai-minimax-h3-en\/\">MiniMax-H3 guide (3 articles)<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-minimax-en\/\">MiniMax: models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-25<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>other<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Unverified (not confirmed by a primary source)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>It is reported that Akatz Labs has released &#8220;akatz-ai\/MiniMax-H3-Character-Swap-LoRA&#8221;, a character swap LoRA adapter designed for the &#8220;MiniMax H3&#8221; video generation model, on Hugging Face. Please note that this information is unverified and has not been officially confirmed.<\/p>\n<p>This model is described as an experimental LoRA that takes an original video (Video-to-Video \/ Ref2VA) as input and uses a reference image to swap only a specific character while preserving the background, lighting, and camera work of the original scene. Only the final checkpoint file at 1,000 steps (1,000 updates) is provided. It does not operate as a standalone model and is distributed as a model-specific LoRA adapter to be applied to the base model with a strength of 1.0.<\/p>\n<h2>Specifications<\/h2>\n<p>Specifications and training configurations confirmed from the model card are as follows:<\/p>\n<ul>\n<li>Base model: <code>Comfy-Org\/MiniMax-H3<\/code> (<code>minimax_h3_ref2va_pruned_int8_convrot.safetensors<\/code>) and <code>ostris\/minimax_h3_training_adapter<\/code> (assistant and base weights are not included in this adapter)<\/li>\n<li>Adapter format: LoRA (not a standalone model or Turbo distillation LoRA)<\/li>\n<li>Training steps: 1,000 updates (checkpoints saved every 250 updates)<\/li>\n<li>LoRA rank \/ alpha: 16 \/ 16 (excluding <code>adaln_proj<\/code>)<\/li>\n<li>Optimizer \/ learning rate: AdamW8bit \/ 5e-5<\/li>\n<li>Precision records: BF16, convrot8 transformer, NVFP4 text encoder<\/li>\n<li>Batch size \/ accumulation: 1 \/ 1<\/li>\n<li>Resolution and area budget: Edit target area budget 1024 (1344\u00d7768 buckets), video regularization area budget 384<\/li>\n<li>Regularization video specs: 73 frames, 24 fps (approx. 3.04 seconds)<\/li>\n<li>Training hardware: RunPod RTX PRO 4500 Blackwell, 32 GB<\/li>\n<li>Recommended settings: LoRA strength 1.0, no trigger word required (describe the target person and swap instructions in the prompt)<\/li>\n<li>License: other<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>Regarding the training settings and records of this model, the publishing model card includes the following table:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Setting<\/th>\n<th>Recorded value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Updates \/ saves<\/td>\n<td>1,000 \/ every 250 updates<\/td>\n<\/tr>\n<tr>\n<td>Hardware<\/td>\n<td>RunPod RTX PRO 4500 Blackwell, 32 GB<\/td>\n<\/tr>\n<tr>\n<td>LoRA rank \/ alpha<\/td>\n<td>16 \/ 16, excluding <code>adaln_proj<\/code><\/td>\n<\/tr>\n<tr>\n<td>Optimizer \/ learning rate<\/td>\n<td>AdamW8bit \/ 5e-5<\/td>\n<\/tr>\n<tr>\n<td>Batch \/ accumulation<\/td>\n<td>1 \/ 1<\/td>\n<\/tr>\n<tr>\n<td>Precision<\/td>\n<td>BF16, convrot8 transformer, NVFP4 text encoder<\/td>\n<\/tr>\n<tr>\n<td>Edit target resolution<\/td>\n<td>Area budget 1024; 1344\u00d7768 buckets<\/td>\n<\/tr>\n<tr>\n<td>Video regularization resolution<\/td>\n<td>Reduced area budget 384<\/td>\n<\/tr>\n<tr>\n<td>Regularization duration<\/td>\n<td>73 frames at 24 fps, approximately 3.04 seconds<\/td>\n<\/tr>\n<tr>\n<td>Memory measures<\/td>\n<td>Gradient checkpointing, layer offload, cached latents\/text, chunked MLP<\/td>\n<\/tr>\n<tr>\n<td>Sampling during training<\/td>\n<td>Disabled<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>According to a local side-by-side evaluation conducted by the publisher, the model excels at preserving the scene and background structure of the original video more faithfully compared to generation using the base model alone. It is reported to have shown promising results, particularly in short, continuous shots of about 4 to 5 seconds.<\/p>\n<p>On the other hand, it is noted that as the generation time increases, composition, placement, and drift from the timing of the original video tend to occur, and clear hard cuts gradually turn into moving zooms or position changes. Furthermore, facial expressions during close-ups may not match the original acting, and it is published that simultaneous swapping of multiple characters is not guaranteed in accuracy since it was not supervised in the training data. Regarding audio generation, continuous tests have also confirmed instances of audio skipping and positional misalignment.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>This model is developed specifically for the character swapping task in Video-to-Video (Ref2VA), where a specific person in the input video is replaced with a character from a reference image or character sheet. It is suitable for applications that apply the appearance, costume, and art style of the character in the reference image to match the movement, pose, and positional relationship of the designated target person, while maintaining the background, lighting, camera work, other characters, props, and environment of the original video as much as possible.<\/p>\n<p>It supports cross-style replacement (such as live-action style to anime style) and swapping from various formats of character sheets. It is said to be particularly easy to obtain good conversion results in short, continuous shots of about 4 to 5 seconds without camera cuts.<\/p>\n<p>On the other hand, processing multiple characters simultaneously in a single generation is not recommended as it was not supervised as a training target. In addition, users need to keep in mind that there are limitations for use cases involving videos with intense cut transitions or those requiring the complete reproduction of delicate mouth movements and facial expressions.<\/p>\n<h2>How It Differences from Similar Models<\/h2>\n<p>The <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/26\/minimax-h3\/\">&#8220;MiniMax-H3&#8221; Video Generation Model with Audio: Required VRAM 32GB+ &amp; File List<\/a> covered previously on our site is a general-purpose omni-modal foundational model that simultaneously generates high-definition video and audio from text, image, and audio inputs, functioning as a complete standalone system. Additionally, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/07\/warmbloodaban-minimax-h3-singularity\/\">&#8220;Minimax-h3_Singularity&#8221; Video Generation Model Running in ComfyUI: Required VRAM 16GB+<\/a> is a fine-tuned model that integrates and optimizes various MiniMax-H3 workflows (T2V, I2V, Ref2V, V2V) for ComfyUI.<\/p>\n<p>In contrast, &#8220;akatz-ai\/MiniMax-H3-Character-Swap-LoRA&#8221; reported this time is decisively different in that it is not a standalone generative model, but an additional adapter (LoRA) applied on top of the base model <code>Comfy-Org\/MiniMax-H3<\/code> to add specific person-swapping capabilities. While it features improved background and scene preservation performance compared to the base model&#8217;s standalone Ref2VA function, it is positioned as an experimental 1,000-step tool specialized in character replacement.<\/p>\n<p><!-- lmw:files --><\/p>\n<h2>Distributed Files<\/h2>\n<p><em>Weight files published in <a href=\"https:\/\/huggingface.co\/akatz-ai\/MiniMax-H3-Character-Swap-LoRA\/tree\/main\">akatz-ai\/MiniMax-H3-Character-Swap-LoRA<\/a>, listed by this site from the Hugging Face API. Sizes are the actual file sizes.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>File<\/th>\n<th>Size<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>h3_character_swap_pro4500_1000.safetensors<\/code><\/td>\n<td>155MB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:files --><\/p>\n<h2>How to Get It<\/h2>\n<p>This LoRA model is distributed in <code>safetensors<\/code> format (<code>h3_character_swap_pro4500_1000.safetensors<\/code>) and can be downloaded directly from the Hugging Face repository <code>akatz-ai\/MiniMax-H3-Character-Swap-LoRA<\/code>. Gated access requirements are not set.<\/p>\n<p>Usage requires an execution environment such as ComfyUI that supports the Ref2VA of MiniMax H3. Place the acquired <code>.safetensors<\/code> file into <code>ComfyUI\/models\/loras\/<\/code> and apply it to the base model using a compatible model-specific LoRA loader. A LoRA strength of 1.0 is recommended as the standard setting during use. Since the base model itself (<code>minimax_h3_ref2va_pruned_int8_convrot.safetensors<\/code>, etc.), VAE, and training assistant weights are not included in this adapter, you must prepare and arrange them separately.<\/p>\n<p>In actual workflows, the original input video and the reference image (or character sheet) for swapping are loaded, and generation is performed by describing the person to be replaced and the background, lighting, and camera actions to be preserved within the prompt. Model-specific trigger words are not trained and are considered unnecessary to specify.<\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p>The publisher distributes this model as safetensors.<\/p>\n<p><strong>License \u2014 <code>other<\/code>:<\/strong> A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats and the license field. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/minimax-h3-open-omnimodal-video-generation\/\">MiniMax-H3 Audio-Visual Video Generation Model: 32GB+ VRAM, File List<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/comfyui-v0350-released\/\">ComfyUI v0.35.0 Released: New Models and Memory Improvements<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Explore the same model family<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-minimaxai-minimax-h3-en\/\">MiniMax-H3 family overview (3 articles, 1 converted builds)<\/a><\/li>\n<li><strong>Formats this model is available in<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">Safetensors format guide and models<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-minimax-en\/\">MiniMax: models, licenses and articles<\/a><\/li>\n<li><strong>Other models for the same task<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-video\">Other video generation models<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/akatz-ai\/MiniMax-H3-Character-Swap-LoRA\">https:\/\/huggingface.co\/akatz-ai\/MiniMax-H3-Character-Swap-LoRA<\/a><\/li>\n<\/ul>\n<blockquote>\n<p><strong>This article contains unverified information.<\/strong> We will append an update note once it is confirmed by a primary source.<\/p>\n<\/blockquote>\n","protected":false},"excerpt":{"rendered":"<p>Discover akatz-ai\/MiniMax-H3-Character-Swap-LoRA, an experimental LoRA adapter for the MiniMax H3 video generation model.<\/p>\n","protected":false},"author":1,"featured_media":6106,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[2534,677,111,730,1565,2387,959],"class_list":["post-6107","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-akatz-ai-en","tag-comfyui-en","tag-lora-en","tag-minimax-h3-en","tag-unverified","tag-video-to-video-en","tag--en"],"lang":"en","translations":{"en":6107,"ja":6105},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6107","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=6107"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6107\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/6106"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=6107"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=6107"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=6107"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}