MiniMax-H3-Character-Swap-LoRA Character Swap LoRA Adapter: File List

MiniMax-H3-Character-Swap-LoRA Character Swap LoRA Adapter: File List

At a Glance

Item Value
Repository akatz-ai/MiniMax-H3-Character-Swap-LoRA
Family guide MiniMax-H3 guide (3 articles)
Publisher guide MiniMax: models and licenses
Published 2026-09-25
License other
Formats safetensors
Source type Unverified (not confirmed by a primary source)

Values determined by this site’s code at collection time. Dates are JST.

Overview

It is reported that Akatz Labs has released “akatz-ai/MiniMax-H3-Character-Swap-LoRA", a character swap LoRA adapter designed for the “MiniMax H3" video generation model, on Hugging Face. Please note that this information is unverified and has not been officially confirmed.

This model is described as an experimental LoRA that takes an original video (Video-to-Video / Ref2VA) as input and uses a reference image to swap only a specific character while preserving the background, lighting, and camera work of the original scene. Only the final checkpoint file at 1,000 steps (1,000 updates) is provided. It does not operate as a standalone model and is distributed as a model-specific LoRA adapter to be applied to the base model with a strength of 1.0.

Specifications

Specifications and training configurations confirmed from the model card are as follows:

  • Base model: Comfy-Org/MiniMax-H3 (minimax_h3_ref2va_pruned_int8_convrot.safetensors) and ostris/minimax_h3_training_adapter (assistant and base weights are not included in this adapter)
  • Adapter format: LoRA (not a standalone model or Turbo distillation LoRA)
  • Training steps: 1,000 updates (checkpoints saved every 250 updates)
  • LoRA rank / alpha: 16 / 16 (excluding adaln_proj)
  • Optimizer / learning rate: AdamW8bit / 5e-5
  • Precision records: BF16, convrot8 transformer, NVFP4 text encoder
  • Batch size / accumulation: 1 / 1
  • Resolution and area budget: Edit target area budget 1024 (1344×768 buckets), video regularization area budget 384
  • Regularization video specs: 73 frames, 24 fps (approx. 3.04 seconds)
  • Training hardware: RunPod RTX PRO 4500 Blackwell, 32 GB
  • Recommended settings: LoRA strength 1.0, no trigger word required (describe the target person and swap instructions in the prompt)
  • License: other

Performance

Regarding the training settings and records of this model, the publishing model card includes the following table:

Setting Recorded value
Updates / saves 1,000 / every 250 updates
Hardware RunPod RTX PRO 4500 Blackwell, 32 GB
LoRA rank / alpha 16 / 16, excluding adaln_proj
Optimizer / learning rate AdamW8bit / 5e-5
Batch / accumulation 1 / 1
Precision BF16, convrot8 transformer, NVFP4 text encoder
Edit target resolution Area budget 1024; 1344×768 buckets
Video regularization resolution Reduced area budget 384
Regularization duration 73 frames at 24 fps, approximately 3.04 seconds
Memory measures Gradient checkpointing, layer offload, cached latents/text, chunked MLP
Sampling during training Disabled

According to a local side-by-side evaluation conducted by the publisher, the model excels at preserving the scene and background structure of the original video more faithfully compared to generation using the base model alone. It is reported to have shown promising results, particularly in short, continuous shots of about 4 to 5 seconds.

On the other hand, it is noted that as the generation time increases, composition, placement, and drift from the timing of the original video tend to occur, and clear hard cuts gradually turn into moving zooms or position changes. Furthermore, facial expressions during close-ups may not match the original acting, and it is published that simultaneous swapping of multiple characters is not guaranteed in accuracy since it was not supervised in the training data. Regarding audio generation, continuous tests have also confirmed instances of audio skipping and positional misalignment.

Strengths and Use Cases

This model is developed specifically for the character swapping task in Video-to-Video (Ref2VA), where a specific person in the input video is replaced with a character from a reference image or character sheet. It is suitable for applications that apply the appearance, costume, and art style of the character in the reference image to match the movement, pose, and positional relationship of the designated target person, while maintaining the background, lighting, camera work, other characters, props, and environment of the original video as much as possible.

It supports cross-style replacement (such as live-action style to anime style) and swapping from various formats of character sheets. It is said to be particularly easy to obtain good conversion results in short, continuous shots of about 4 to 5 seconds without camera cuts.

On the other hand, processing multiple characters simultaneously in a single generation is not recommended as it was not supervised as a training target. In addition, users need to keep in mind that there are limitations for use cases involving videos with intense cut transitions or those requiring the complete reproduction of delicate mouth movements and facial expressions.

How It Differences from Similar Models

The “MiniMax-H3" Video Generation Model with Audio: Required VRAM 32GB+ & File List covered previously on our site is a general-purpose omni-modal foundational model that simultaneously generates high-definition video and audio from text, image, and audio inputs, functioning as a complete standalone system. Additionally, “Minimax-h3_Singularity" Video Generation Model Running in ComfyUI: Required VRAM 16GB+ is a fine-tuned model that integrates and optimizes various MiniMax-H3 workflows (T2V, I2V, Ref2V, V2V) for ComfyUI.

In contrast, “akatz-ai/MiniMax-H3-Character-Swap-LoRA" reported this time is decisively different in that it is not a standalone generative model, but an additional adapter (LoRA) applied on top of the base model Comfy-Org/MiniMax-H3 to add specific person-swapping capabilities. While it features improved background and scene preservation performance compared to the base model’s standalone Ref2VA function, it is positioned as an experimental 1,000-step tool specialized in character replacement.

Distributed Files

Weight files published in akatz-ai/MiniMax-H3-Character-Swap-LoRA, listed by this site from the Hugging Face API. Sizes are the actual file sizes.

File Size
h3_character_swap_pro4500_1000.safetensors 155MB

How to Get It

This LoRA model is distributed in safetensors format (h3_character_swap_pro4500_1000.safetensors) and can be downloaded directly from the Hugging Face repository akatz-ai/MiniMax-H3-Character-Swap-LoRA. Gated access requirements are not set.

Usage requires an execution environment such as ComfyUI that supports the Ref2VA of MiniMax H3. Place the acquired .safetensors file into ComfyUI/models/loras/ and apply it to the base model using a compatible model-specific LoRA loader. A LoRA strength of 1.0 is recommended as the standard setting during use. Since the base model itself (minimax_h3_ref2va_pruned_int8_convrot.safetensors, etc.), VAE, and training assistant weights are not included in this adapter, you must prepare and arrange them separately.

In actual workflows, the original input video and the reference image (or character sheet) for swapping are loaded, and generation is performed by describing the person to be replaced and the background, lighting, and camera actions to be preserved within the prompt. Model-specific trigger words are not trained and are considered unnecessary to specify.

Can You Run It Locally?

The publisher distributes this model as safetensors.

License — other: A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.

Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Related Articles

What to Read Next

Sources

This article contains unverified information. We will append an update note once it is confirmed by a primary source.