Replace the person in an existing video with a new character defined by one reference image - while the original scene, camera movement, lighting and the original audio track stay untouched.
This workflow is the single-video depth composite design for MiniMax H3. Video 1 and Picture 1 are the only two inputs: no second video, no first/last frame pair.
How the chain works
- Video 1 is loaded at 24 fps (124-frame cap in the shipped config) and downscaled to ~0.94 MP on 32-pixel steps (sample: 1152x864, 124 frames).
- You mark positive points on the target person and negative points on exclusions in the PointsEditor. SeC (SeCVideoSegmentation) tracks the person across the whole video; GrowMask (+8 px) expands the person region.
- Video Depth Anything (ViT-S) extracts full-frame grayscale depth for every frame.
- ImageCompositeMasked keeps the original RGB everywhere except the tracked person, which is replaced by its depth rendering. The result is ONE hybrid Video 1: RGB scene + Depth person.
- MiniMaxH3ReferenceToVideo (Ref2VA) receives Picture 1 as the identity reference and the hybrid video as the structural reference. The built-in subject prompt defines three subjects: the replacement character from Picture 1, the RGB regions to preserve, and the depth region that only guides pose, silhouette, 3D volume, limb ordering and motion.
- The original audio is encoded with the H3 audio VAE and locked (SolidMask + SetLatentNoiseMask) so the soundtrack passes through without regeneration.
- Sampling: the MiniMax H3 Hybrid INT8 base is composed with OpenVDN (vdn-minimax-h3, stage_dmd_8nfe), which runs the fixed 8-NFE execution plan in SamplerCustomAdvanced (fixed seed 9527), then video VAE decode and CreateVideo mux in the original audio before SaveVideo.
Main features:
- One video + one image input; single-video design, no second clip required
- Original scene preservation: all unmasked RGB content (background, framing, camera movement, lighting, other people, props) is kept
- Original audio passthrough: the audio latent is locked and never re-synthesized
- Fast generation: OpenVDN DMD8 distillation stage, fixed 8 NFE
- Structural control: the depth region drives body position, pose, limb depth ordering, gesture trajectory and spatial placement; it is never rendered as gray texture
- No LoRA needed - character identity comes entirely from the Picture 1 reference
Suggested workflow:
1. Load only Video 1 (the clip containing the person to replace) and Picture 1 (the new character reference).
2. In the PointsEditor, mark positive points on the target person and negative points on things that must not be included; let SeC track the whole video.
3. Inspect the tracked mask before running. GrowMask defaults to 8 - adjust it so the background around the person is not converted to depth.
4. Queue the prompt. The output MP4 keeps the scene and the original audio while the person is replaced with the Picture 1 identity.
Model and node dependencies:
- UNet: minimax_h3_hybrid_fl2va_ref2va_zs05_b25-49_int8.safetensors (MiniMax H3 Hybrid INT8, OpenVDN structural base)
- CLIP: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (Qwen3VL 32B NVFP4, type: minimax)
- Video VAE: minimax_h3_video_vae_int8_convrot.safetensors
- Audio VAE: minimax_h3_audio_vae_fp32.safetensors
- Depth: video_depth_anything_vits.pth (Video Depth Anything ViT-S)
- SeC model loaded via SeCModelLoader (bfloat16 / auto device)
- Custom nodes: minimax-h3-audio-T8 (MiniMaxH3ReferenceToVideo, MiniMaxH3VDNModelComposerT8Advanced, MiniMaxH3VDNExecutionPlanT8Advanced, LTXVAudioVAEEncode, LTXVConcatAVLatent, LTXVSeparateAVLatent), ComfyUI-VideoHelperSuite (VHS_LoadVideo), ComfyUI-VideoDepthAnything (LoadVideoDepthAnythingModel, VideoDepthAnythingProcess), SeC nodes (SeCModelLoader, SeCVideoSegmentation), KJNodes (PointsEditor)
Inputs / outputs:
- Inputs: 1 video (MP4, up to 124 frames in the shipped config) + 1 character reference image
- Output: MP4 at ~0.94 MP (sample 1152x864) at the source frame rate (24 fps), original audio track preserved
Limitations:
- Designed for one main person per video; multiple people require separate masked passes
- The shipped config caps at 124 frames - raise frame_load_cap for longer clips (VRAM dependent)
- Fixed 8 NFE (DMD8 stage): fast, distilled quality
- The grayscale depth area is control information only - the prompt forbids it from appearing as visible gray
Bypassed / unused nodes: none - all 31 nodes are active in this workflow.
Links (to be added as they go live - none provided for this batch, none fabricated):
- YouTube: pending
- Bilibili: pending
- RunningHub: pending
This workflow is the single-video depth composite design for MiniMax H3. Video 1 and Picture 1 are the only two inputs: no second video, no first/last frame pair.
How the chain works
- Video 1 is loaded at 24 fps (124-frame cap in the shipped config) and downscaled to ~0.94 MP on 32-pixel steps (sample: 1152x864, 124 frames).
- You mark positive points on the target person and negative points on exclusions in the PointsEditor. SeC (SeCVideoSegmentation) tracks the person across the whole video; GrowMask (+8 px) expands the person region.
- Video Depth Anything (ViT-S) extracts full-frame grayscale depth for every frame.
- ImageCompositeMasked keeps the original RGB everywhere except the tracked person, which is replaced by its depth rendering. The result is ONE hybrid Video 1: RGB scene + Depth person.
- MiniMaxH3ReferenceToVideo (Ref2VA) receives Picture 1 as the identity reference and the hybrid video as the structural reference. The built-in subject prompt defines three subjects: the replacement character from Picture 1, the RGB regions to preserve, and the depth region that only guides pose, silhouette, 3D volume, limb ordering and motion.
- The original audio is encoded with the H3 audio VAE and locked (SolidMask + SetLatentNoiseMask) so the soundtrack passes through without regeneration.
- Sampling: the MiniMax H3 Hybrid INT8 base is composed with OpenVDN (vdn-minimax-h3, stage_dmd_8nfe), which runs the fixed 8-NFE execution plan in SamplerCustomAdvanced (fixed seed 9527), then video VAE decode and CreateVideo mux in the original audio before SaveVideo.
Main features:
- One video + one image input; single-video design, no second clip required
- Original scene preservation: all unmasked RGB content (background, framing, camera movement, lighting, other people, props) is kept
- Original audio passthrough: the audio latent is locked and never re-synthesized
- Fast generation: OpenVDN DMD8 distillation stage, fixed 8 NFE
- Structural control: the depth region drives body position, pose, limb depth ordering, gesture trajectory and spatial placement; it is never rendered as gray texture
- No LoRA needed - character identity comes entirely from the Picture 1 reference
Suggested workflow:
1. Load only Video 1 (the clip containing the person to replace) and Picture 1 (the new character reference).
2. In the PointsEditor, mark positive points on the target person and negative points on things that must not be included; let SeC track the whole video.
3. Inspect the tracked mask before running. GrowMask defaults to 8 - adjust it so the background around the person is not converted to depth.
4. Queue the prompt. The output MP4 keeps the scene and the original audio while the person is replaced with the Picture 1 identity.
Model and node dependencies:
- UNet: minimax_h3_hybrid_fl2va_ref2va_zs05_b25-49_int8.safetensors (MiniMax H3 Hybrid INT8, OpenVDN structural base)
- CLIP: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (Qwen3VL 32B NVFP4, type: minimax)
- Video VAE: minimax_h3_video_vae_int8_convrot.safetensors
- Audio VAE: minimax_h3_audio_vae_fp32.safetensors
- Depth: video_depth_anything_vits.pth (Video Depth Anything ViT-S)
- SeC model loaded via SeCModelLoader (bfloat16 / auto device)
- Custom nodes: minimax-h3-audio-T8 (MiniMaxH3ReferenceToVideo, MiniMaxH3VDNModelComposerT8Advanced, MiniMaxH3VDNExecutionPlanT8Advanced, LTXVAudioVAEEncode, LTXVConcatAVLatent, LTXVSeparateAVLatent), ComfyUI-VideoHelperSuite (VHS_LoadVideo), ComfyUI-VideoDepthAnything (LoadVideoDepthAnythingModel, VideoDepthAnythingProcess), SeC nodes (SeCModelLoader, SeCVideoSegmentation), KJNodes (PointsEditor)
Inputs / outputs:
- Inputs: 1 video (MP4, up to 124 frames in the shipped config) + 1 character reference image
- Output: MP4 at ~0.94 MP (sample 1152x864) at the source frame rate (24 fps), original audio track preserved
Limitations:
- Designed for one main person per video; multiple people require separate masked passes
- The shipped config caps at 124 frames - raise frame_load_cap for longer clips (VRAM dependent)
- Fixed 8 NFE (DMD8 stage): fast, distilled quality
- The grayscale depth area is control information only - the prompt forbids it from appearing as visible gray
Bypassed / unused nodes: none - all 31 nodes are active in this workflow.
Links (to be added as they go live - none provided for this batch, none fabricated):
- YouTube: pending
- Bilibili: pending
- RunningHub: pending
Description
Replace the person in an existing video with a new character defined by one reference image - while the original scene, camera movement, lighting and the original audio track stay untouched.
This workflow is the single-video depth composite design for MiniMax H3. Video 1 and Picture 1 are the only two inputs: no second video, no first/last frame pair.
How the chain works
- Video 1 is loaded at 24 fps (124-frame cap in the shipped config) and downscaled to ~0.94 MP on 32-pixel steps (sample: 1152x864, 124 frames).
- You mark positive points on the target person and negative points on exclusions in the PointsEditor. SeC (SeCVideoSegmentation) tracks the person across the whole video; GrowMask (+8 px) expands the person region.
- Video Depth Anything (ViT-S) extracts full-frame grayscale depth for every frame.
- ImageCompositeMasked keeps the original RGB everywhere except the tracked person, which is replaced by its depth rendering. The result is ONE hybrid Video 1: RGB scene + Depth person.
- MiniMaxH3ReferenceToVideo (Ref2VA) receives Picture 1 as the identity reference and the hybrid video as the structural reference. The built-in subject prompt defines three subjects: the replacement character from Picture 1, the RGB regions to preserve, and the depth region that only guides pose, silhouette, 3D volume, limb ordering and motion.
- The original audio is encoded with the H3 audio VAE and locked (SolidMask + SetLatentNoiseMask) so the soundtrack passes through without regeneration.
- Sampling: the MiniMax H3 Hybrid INT8 base is composed with OpenVDN (vdn-minimax-h3, stage_dmd_8nfe), which runs the fixed 8-NFE execution plan in SamplerCustomAdvanced (fixed seed 9527), then video VAE decode and CreateVideo mux in the original audio before SaveVideo.
Main features:
- One video + one image input; single-video design, no second clip required
- Original scene preservation: all unmasked RGB content (background, framing, camera movement, lighting, other people, props) is kept
- Original audio passthrough: the audio latent is locked and never re-synthesized
- Fast generation: OpenVDN DMD8 distillation stage, fixed 8 NFE
- Structural control: the depth region drives body position, pose, limb depth ordering, gesture trajectory and spatial placement; it is never rendered as gray texture
- No LoRA needed - character identity comes entirely from the Picture 1 reference
Suggested workflow:
1. Load only Video 1 (the clip containing the person to replace) and Picture 1 (the new character reference).
2. In the PointsEditor, mark positive points on the target person and negative points on things that must not be included; let SeC track the whole video.
3. Inspect the tracked mask before running. GrowMask defaults to 8 - adjust it so the background around the person is not converted to depth.
4. Queue the prompt. The output MP4 keeps the scene and the original audio while the person is replaced with the Picture 1 identity.
Model and node dependencies:
- UNet: minimax_h3_hybrid_fl2va_ref2va_zs05_b25-49_int8.safetensors (MiniMax H3 Hybrid INT8, OpenVDN structural base)
- CLIP: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (Qwen3VL 32B NVFP4, type: minimax)
- Video VAE: minimax_h3_video_vae_int8_convrot.safetensors
- Audio VAE: minimax_h3_audio_vae_fp32.safetensors
- Depth: video_depth_anything_vits.pth (Video Depth Anything ViT-S)
- SeC model loaded via SeCModelLoader (bfloat16 / auto device)
- Custom nodes: minimax-h3-audio-T8 (MiniMaxH3ReferenceToVideo, MiniMaxH3VDNModelComposerT8Advanced, MiniMaxH3VDNExecutionPlanT8Advanced, LTXVAudioVAEEncode, LTXVConcatAVLatent, LTXVSeparateAVLatent), ComfyUI-VideoHelperSuite (VHS_LoadVideo), ComfyUI-VideoDepthAnything (LoadVideoDepthAnythingModel, VideoDepthAnythingProcess), SeC nodes (SeCModelLoader, SeCVideoSegmentation), KJNodes (PointsEditor)
Inputs / outputs:
- Inputs: 1 video (MP4, up to 124 frames in the shipped config) + 1 character reference image
- Output: MP4 at ~0.94 MP (sample 1152x864) at the source frame rate (24 fps), original audio track preserved
Limitations:
- Designed for one main person per video; multiple people require separate masked passes
- The shipped config caps at 124 frames - raise frame_load_cap for longer clips (VRAM dependent)
- Fixed 8 NFE (DMD8 stage): fast, distilled quality
- The grayscale depth area is control information only - the prompt forbids it from appearing as visible gray
Bypassed / unused nodes: none - all 31 nodes are active in this workflow.
Links (to be added as they go live - none provided for this batch, none fabricated):
- YouTube: pending
- Bilibili: pending
- RunningHub: pending
This workflow is the single-video depth composite design for MiniMax H3. Video 1 and Picture 1 are the only two inputs: no second video, no first/last frame pair.
How the chain works
- Video 1 is loaded at 24 fps (124-frame cap in the shipped config) and downscaled to ~0.94 MP on 32-pixel steps (sample: 1152x864, 124 frames).
- You mark positive points on the target person and negative points on exclusions in the PointsEditor. SeC (SeCVideoSegmentation) tracks the person across the whole video; GrowMask (+8 px) expands the person region.
- Video Depth Anything (ViT-S) extracts full-frame grayscale depth for every frame.
- ImageCompositeMasked keeps the original RGB everywhere except the tracked person, which is replaced by its depth rendering. The result is ONE hybrid Video 1: RGB scene + Depth person.
- MiniMaxH3ReferenceToVideo (Ref2VA) receives Picture 1 as the identity reference and the hybrid video as the structural reference. The built-in subject prompt defines three subjects: the replacement character from Picture 1, the RGB regions to preserve, and the depth region that only guides pose, silhouette, 3D volume, limb ordering and motion.
- The original audio is encoded with the H3 audio VAE and locked (SolidMask + SetLatentNoiseMask) so the soundtrack passes through without regeneration.
- Sampling: the MiniMax H3 Hybrid INT8 base is composed with OpenVDN (vdn-minimax-h3, stage_dmd_8nfe), which runs the fixed 8-NFE execution plan in SamplerCustomAdvanced (fixed seed 9527), then video VAE decode and CreateVideo mux in the original audio before SaveVideo.
Main features:
- One video + one image input; single-video design, no second clip required
- Original scene preservation: all unmasked RGB content (background, framing, camera movement, lighting, other people, props) is kept
- Original audio passthrough: the audio latent is locked and never re-synthesized
- Fast generation: OpenVDN DMD8 distillation stage, fixed 8 NFE
- Structural control: the depth region drives body position, pose, limb depth ordering, gesture trajectory and spatial placement; it is never rendered as gray texture
- No LoRA needed - character identity comes entirely from the Picture 1 reference
Suggested workflow:
1. Load only Video 1 (the clip containing the person to replace) and Picture 1 (the new character reference).
2. In the PointsEditor, mark positive points on the target person and negative points on things that must not be included; let SeC track the whole video.
3. Inspect the tracked mask before running. GrowMask defaults to 8 - adjust it so the background around the person is not converted to depth.
4. Queue the prompt. The output MP4 keeps the scene and the original audio while the person is replaced with the Picture 1 identity.
Model and node dependencies:
- UNet: minimax_h3_hybrid_fl2va_ref2va_zs05_b25-49_int8.safetensors (MiniMax H3 Hybrid INT8, OpenVDN structural base)
- CLIP: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (Qwen3VL 32B NVFP4, type: minimax)
- Video VAE: minimax_h3_video_vae_int8_convrot.safetensors
- Audio VAE: minimax_h3_audio_vae_fp32.safetensors
- Depth: video_depth_anything_vits.pth (Video Depth Anything ViT-S)
- SeC model loaded via SeCModelLoader (bfloat16 / auto device)
- Custom nodes: minimax-h3-audio-T8 (MiniMaxH3ReferenceToVideo, MiniMaxH3VDNModelComposerT8Advanced, MiniMaxH3VDNExecutionPlanT8Advanced, LTXVAudioVAEEncode, LTXVConcatAVLatent, LTXVSeparateAVLatent), ComfyUI-VideoHelperSuite (VHS_LoadVideo), ComfyUI-VideoDepthAnything (LoadVideoDepthAnythingModel, VideoDepthAnythingProcess), SeC nodes (SeCModelLoader, SeCVideoSegmentation), KJNodes (PointsEditor)
Inputs / outputs:
- Inputs: 1 video (MP4, up to 124 frames in the shipped config) + 1 character reference image
- Output: MP4 at ~0.94 MP (sample 1152x864) at the source frame rate (24 fps), original audio track preserved
Limitations:
- Designed for one main person per video; multiple people require separate masked passes
- The shipped config caps at 124 frames - raise frame_load_cap for longer clips (VRAM dependent)
- Fixed 8 NFE (DMD8 stage): fast, distilled quality
- The grayscale depth area is control information only - the prompt forbids it from appearing as visible gray
Bypassed / unused nodes: none - all 31 nodes are active in this workflow.
Links (to be added as they go live - none provided for this batch, none fabricated):
- YouTube: pending
- Bilibili: pending
- RunningHub: pending
comfyui
workflow
workflows
video editing
character replacement
minimax h3
ref2va
dmd8
openvdn
video depth anything
single video
depth composite
Details
Downloads
81
Platform
CivitAI
Platform Status
Available
Created
10/1/2026
Updated
10/2/2026
Deleted
-
