CivArchive
    Minimax H3 Faceswap with Spectrum & Sparse Attention - v1.0
    NSFW
    Preview 144200899

    This workflow performs video face swaps on my system with a 100% success rate. Contains some custom nodes that are available for install via Comfy Manager. If you can't find them, use an LLM to help you - none are hard to get or behind paywalls so you'll install them easy enough.

    Does better with source videos that aren't completely side-on for the duration but handles hands/other objects obscuring the face with no difficulty.

    INSTRUCTIONS:

    Leave the SAM3 prompt as it is (or adjust to 'male's face' if that's what you're into).

    Load your clipped action video into the Load Video node (top left).

    Load your front, three-quarter, side and 'extra' closeup images into the four Load Image nodes at the bottom.

    LEAVE THE ENABLE LIGHTNING SWITCH UNDER THE MODELS SECTION SWITCHED TO 'TRUE'!

    Resolution: currently set to 1344 x 768, native MMH3 resolution. If you change it, you HAVE to also change the Resize Image/Mask node in the Masking section so that they match - if you don't do this the masking will not work.

    Leave all other settings as they are, unless you know exactly what you're doing and why.

    The AV latent from the masking is what the KSampler uses, this keeps the action and sound from the source video intact. The Minimax H3 node provides the guidance to make the facial magic happen.

    Source clips should be at least 5 seconds duration. They can be longer, the Math Expression takes care of that so always leave the Float (Duration) value set to 5. However long your clip is, the generated output will match it thanks to the AV latent.

    Please use responsibly and legally.

    Generates 10 second clips in under 10 minutes on my setup (9070XT, 16GB VRAM, 64GB RAM) so if you've got a decent NVIDIA card it should be way quicker. It's optimised for AMD though so you may need to tweak things, try a run and see.

    Models can be changed but I've been running on:

    Minimax H3 ref2va pruned INT8 convrot - Diffusion Model

    Minimax H3 ref2v turbo 4step v0.1 comfyui bf16 - LoRA

    Qwen3vl 4b fp8 scaled - CLIP (into MM3 4b ClipProj v3.1 mlp, which can be removed but generations will take longer)

    Minimax H3 video vae fp16 - Video VAE

    Minimax H3 audio vae fp32 - Audio VAE

    Comfy version I've used: 0.35.1

    Startup arguments: --enable-manager --fast-disk --use-ck-attention --fast (I'm using an AMD 9070XT so YMMV)

    OS environment variables (again, YMMV): COMFYUI_ENABLE_MIOPEN=1, HIP_FORCE_DEV_KERNARG=1, MIOPEN_FIND_MODE=FAST, OPTIMIZE_FOR_SPEED=1, PYTORCH_ALLOC_CONF=expandable_segments:True, PYTORCH_TUNABLEOP_ENABLED=1, SAFETENSORS_FAST_GPU=1, TORCH_BLAS_PREFER_HIPBLASLT=1, TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1

    Optional but recommended: PYTORCH_TUNABLEOP_TUNING=1 (on first run, which will make it relatively slow. Then set to 0 for all subsequent runs with the same settings and they'll always be optimised).

    PROMPT (Don't change)

    <Picture 1>, <Picture 2>, <Picture 3> and <Picture 4> all depict the same woman and define one consistent facial identity.

    <Picture 1> is the primary identity reference.

    <Picture 2> provides additional three-quarter facial structure.

    <Picture 3> provides additional side-profile facial structure.

    <Picture 4> provides additional information about her natural facial appearance.

    Use all four pictures together to maintain one consistent identity. Do not mix identities. The person in the generated video must have the facial identity defined by the four picture references.

    Use the reference images only as the identity and facial appearance reference.

    The source video provides the facial performance and motion.

    Preserve the source video's:

    facial expressions,

    eye movements,

    blinking,

    mouth movements,

    lip movements,

    head movement,

    facial motion,

    timing and performance.

    Replace the identity of the person in the masked region with the woman represented by <Picture 1>, <Picture 2>, <Picture 3> and <Picture 4>.

    Do not freeze, neutralize, or simplify the facial expression.

    Do not just reproduce the facial expression from the reference pictures.

    The expression must follow the source video throughout the sequence.

    Keep the target identity consistent across all frames.

    non_diegetic_music: N/A

    Enjoy!

    Description

    See description, I accidentally wrote all the notes in there

    FAQ

    Comments (3)

    jafdeth2030105Sep 30, 2026
    CivitAI

    Friend, I updated my Comfy to 0.38, but the h3 apply mask node still says:
    Exception Message: RuntimeError: h3_masking: general masked-target generation requires the native MiniMax H3 Add Guide / MultiRef core from ComfyUI PR #15439. Update ComfyUI before using this path.

    What else does workflow need?

    unclejunsfartcushion
    Author
    Oct 1, 2026· 1 reaction

    Install via ComfyUI Manager: Open ComfyUI Click Manager (or 'Extensions' in main menu). Go to Custom Nodes Manager / Install Custom Nodes. Search for: MiniMax H3 PerRowMasking Look for: ComfyUI-MiniMaxH3-PerRowMasking. Click Install.

    Restart ComfyUI completely — close the ComfyUI server/terminal and restart it, rather than just refreshing the browser. After restarting, search the node menu for MiniMax H3. You should see these experimental nodes: MiniMax H3 Trim Source AV to 17k+5 MiniMax H3 Per-Row Mask Patch (Experimental) MiniMax H3 Set Generation Mask MiniMax H3 Mask Grid Preview & Snap

    If that doesn't work, it's because you're on 0.38. The repository's documented installation method is currently: ComfyUI/custom_nodes/ComfyUI-MiniMaxH3-PerRowMasking via Git clone, followed by a restart.

    jafdeth2030105Oct 1, 2026

    @unclejunsfartcushion your advice helped - it works. Thanks! 

    Workflows
    MiniMax H3

    Details

    Downloads
    728
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/29/2026
    Updated
    10/6/2026
    Deleted
    -

    Files

    minimaxH3FaceswapWith_v10.json

    Mirrors