This workflow uses SAM to invert the colors of the character to be replaced. By doing so, it makes character replacement significantly more consistent that Minimax H3 alone.
How to use it
You must
(v2 and above) Install https://github.com/kijai/ComfyUI-SolAttn_triton if you want the speedup.
(v3) Install https://github.com/1038lab/ComfyUI-QwenVL for QwenVL video to text description.
(v1 only) Change the text prompt to describe both the scenario and the subject. Your goal is that Minimax H3 reproduces the original video as similar as possible.
Change the video and the picture, and make sure to change as well the resolution to match the original video.
(Optional, v2) change the SAM prompt until it masks only the subject to replace.
(Optional, v3) change the hint given to the Qwen node.
Why SAM and masking with H3
When using H3, it seems that the model will either recreate the original video as it is, without changing the subject at all (follows the <Video 1> too literally), or it will create a totally new video (follows the <Picture 1> too literally).
However, when inverting the color of the subject in the video, H3 is forced to produce a new video because the subject looks weird. So it can't follow the <Video 1> too literally anymore, otherwise it would produce a video with inverted colors.
Description
With QwenVL, describes the scene and the subject to do the replacement
FAQ
Comments (48)
Which version of Qwen VL should we be using for the Qwen-VL Mod node?
I tested with a few options and keep getting the following error:
AILab_QwenVL_Advanced.process() got an unexpected keyword argument 'video'
Specifically the first QwenVL Mod that does the video analysis.
The node I use automatically downloads the model I think. I don't remember downloading it manually.
@lomote The QWENVL model's censored though, it doesnt accurately describe the scene.
@mogmog I'm not sure which VL models are not censured, so feel free to suggest a drop in replacement.
@lomote I think there's a built in node with the same name, because it wasnt orignally flagged as a missing node. I did install that version and tested again with the same error.
I've tried a few different videos and Qwen models from the drop down.
I've rarely had luck with the qwen llm nodes, so maybe its something on my backend (nightly comfy and backend python should be up-to-date via the .bat file), but thats the only error im getting.
I think I figured it out. For some reason the qwenvl mod node was not updating (noticed it only had 2 image inputs when i tried to re-add the node). Blew away both the Comfyui-QwenVL (and Qwen-VL Mod) folders and reinstalled with: pip install --upgrade --force-reinstall <package>. Appears to be running now.
@mogmog you can click on it and change to the abliterated or heretic versions that are not censored
If we didnt want to use the reference audio but generate new audio instead, how would we rewire it?
This has worked really well for face/head swap.
Just change the SAM prompt to "The woman's head". I also recommend adding this enforcement throughout the generation prompt.
Example (insert features/anatomy you want to stay identical):
The target video replaces only the head and face of the woman in <Video 1> with <Subject 1> while fully preserving the entire body below the neck, the exact _______, the low-angle camera, the environment, and the lighting.
Hi is it just me or what background not being copied from video, movemnet in dance is just being "generalized" not copied. Some movemnet are just being ingnored and model replace it with something else. Bad prompting or other issue?
The last example of the dance didn't copy exactly, and there are artifacts. With other simpler scenes this works better.
Here the prompt was produced automatically with QwenVL. I don't know if it can be improved further, but if you get a prompt let me know.
@lomote I think we need something like fine-tune with better conditioning - maybe like SCAIL-2?
great stuff, the only problem right now is liquids hitting the face, with invert the model just makes a blue ink for any white liquids, there has to be a way to fix it but I cannot figure out.
"White liquids"? (╭ರ_•́)
Really like this workflow, especially the SAM invert trick.
You can add KJNodes Model Preview Override for sampling previews, makes it a lot easier to keep an eye on longer H3 runs: https://github.com/kijai/ComfyUI-KJNodes/blob/main/nodes/preview_override_node.py
Also found a nice speed win: added a second Scale Image to Total Pixels right after Image Composite Masked, before the reference video goes into H3, and dropped it from 0.5 to 0.2 megapixels. Sampling time went from 3:01 to 2:01 (roughly 1.5x), total workflow time from 280s to 212s. Compared both outputs 1:1, motion and detail included, barely any difference.
[edit]: Saw you already mentioned this, might be worth just adding the node in for people to tweak as needed.
One more thing that could help: an optional pause after QwenVL and before H3, to check or tweak the prompt before the expensive sampling starts. Would be nice if it's opt-in so it doesn't slow down a normal run.
I'm newish & don't understand; can you provide a workflow please?
for live preview: install the node from the link above and wire it as Load Diffusion Model > Model Preview Override > Turbo Lora.
for the 1.5x speed boost: clone "Scale Image to Total Pixels", place it between Image Composite Masked and MiniMax H3 Reference to Video, and set it to 0.2
It really sucks that the inverting hack needs to be done to reliably swap a character, otherwise the model just completely ignores one's prompting. For a 'state of the art' model that everyone claims this thing is, it seems like kind of a joke that it can't just do this natively. And yes, I know all about how the r2v prompts are supposed to be structured in h3; it usually doesn't make a difference, even with a perfectly structured prompt. Unless the characters you're swapping are literally visual opposites from each other, the model is just going to keep the original subject no matter how clear your prompt is...
Anyway, venting over; thanks for the hack. Glad it's working out for people.
Awesome, what if your SAM3 patch could do a full motion transfer? (Image background stays). Anyways. Left you a large sum - don't spend it all in one place XD :)
Thanks a lot! I'm happy to see people finding this useful.
About your comment, I don't understand what do you mean by full motion. This is what I have done with all the examples, right? The character is replaced but the general movement and scene is preserved. Can you give me an example of what are you missing?
@lomote Awesome. I am saying like full motion transfer like Kling. Not character replacement like your wonderful workflow.
ImportError: finegrained-fp8 kernel requires the kernels package. Install it with pip install -U kernels.
*for people that have this error, change the qwen3 model to one that doesnt have "fp8"
Anyone know how to fix this error? Tried a few things but keep running into it, cant run the v3 version.
# ComfyUI Error Report
## Error Details
- Node ID: 921
- Node Type: AILab_QwenVL_Advanced
- Exception Type: ValueError
- Exception Message: ValueError: Kernel repository 'kernels-community/finegrained-fp8' could not verify publisher trust status. Set trust_remote_code=True to allow loading kernels from untrusted sources.
Ctrl+v it into any free ai model. Error literally says where problem is.
Why free ai and not some gpt or claude (and not clear answer how to fix it)? Because free ai is educated and trained on same datasets as elses, but pretty stupid (after they rush for native cards). So, both of you together (because you are smart, but not experienced) will resolve this problem (maybe not fast, but you will knew something helpful in future terms).
This works very well, could you make one that can incorporate a male swap as well alongside a female swap?
hi trying to make it work but kept meeting issue with Qwen V3, it just appears as a red outline and i am not able to install the node, will i really need to use the Qwen V3 to run this workflow?
No, you can use V2 or V1. If you provide your own description of the scene and the image in V2, it should work even better than V3.
Why not use the Qwen3-VL 32B that's already downloaded and being used for H3, instead of this node that downloads and installs a separate 8B one?
I wasn't aware these were actually the same model. Is there any way to reuse it so we don't have to load it twice?
the 32b model is huge and would take much longer to go through the video. an 8b fp8 model is a fraction of the size and you can use an abliterated model for better uncensored prompting. the node also unloads the model afterwards so it frees up your vram after generating the text. you could even use smaller models for faster inference its up to you what model to use for qwen vl.
i noticed that the pacing of the original video is usually (but not always) lost. For example an 8 second video of 3 things happening becomes an 8 second video with 2 things happening, at a slower pace that is no longer synced to the original audio. has anyone figured this out? I'm not using turbo loras or the patch, both nodes are bypassed. increased steps from 8 to 12
I would say, mismatch of input-output framerate. If the motion is copied frame by frame and the input video has a different fps count, it will result in a different outcome since every frame is compared - and not the duration of the input clip. So if an input clip is at 60 fps - that would mean you have 300 frames in 5 seconds. A 5 second output at 24 frames has 120 frames - in this case, only about the first third of the input video would be in the output and it would also be in slow motion.
@claude111 thank you that makes a lot of sense, probably is my issue.
nice and I've found if you add additional reference image or video (even better) with close-up face shot it gives better face similarity, you just add in the instruct node something like "<Subject 1> is the woman with body in <Picture 1> and her facial features identical to <Video 2>."
Countless node errors occurred
Fellows, the Draw ViT Pose and DWPOse Estimator is works fine too! Like in the Wan Animate!!
How does that work? Would you like to share more?
Works very well. I didn't use your exact workflow but that invert color concept is very clever. nice find
Did you successfully replace the character? The mask thumbnail shows that it is indeed the character I wanted to replace, but the actual video generation fails. The character is replaced, but it feels like the reference image was used to generate the new video. The original video content and mask are not functioning.
@fctq9169 yes, but i made a different wf inspired by this, https://github.com/bitsofintelligence101-lab/workflows/blob/main/nsfw/h3/minimax_h3_sam_r2v_cinematic.json
@bitsofintelligence101907 I see it!!! It seems even more perfect. Currently in testing. Should a timeline description be added to the prompt? If it's a scene where the character is naked, should a description of their body shape be included?
@fctq9169 that stuff all helps. i setup the prompt to be 'omni' so you don't need to change anything except adding spoken dialogue if there is any. there are some notes in the workflow too over to the side about that
I'd like to ask if anyone has successfully used this method. I'm using v20SolAttnTurbo to replace a 5-second video for testing. The sam+mask thumbnail does indeed mask and color the character I specified. However, I don't know if it's a node connection issue or a cue word issue, but the generated video can't match the original video's mask position. The character is successfully replaced in the new video, but it's completely new content. The character's position, size, screen proportion, and even movements are different from the original video. My cue words are the MMH3 SKILL and the video reference image sent to Grok for AI to write. The content seems fine, but it still fails. Does anyone know what the possible cause is?
I'm using V2+turbo. Currently, I'm encountering an issue where landscape videos are almost never successfully generated. The mask thumbnails are successful, and one of the landscape videos works, but it fails with the other reference images. Does anyone know what's causing this? Is it a sampler problem? A seed problem? Or something else?
@lomote you're a genius. Nice work!
Out of curiosity what chip/stack do you recommend running this on? And any tips for improving speed of rendering with little to no loss of quality?
Anyone know a way to "auto" skip the QwenVL node? The SAM3 node is smart enough to not run again on the same video, but the QwenVL is not and it's a waste to do so when the video hasn't changed.
Set a manual multiline text node, and either connect it or manually paste the text in there. That is the easiest way.