Lipsync with Minimax H3 Ref2VA.
There is no transcript or lyrics needed in the prompt. The model is smart enough to lipsync, as long as you prompt it to.
This is the default ComfyUI workflow for Ref2VA, with:
Kijai's LightX2V LoRA and Kijai's sage attention patch enabled
Math to calculate the length of the audio track to match the number of frames
Due to request from community, I made a version with 2 speakers for a podcast. This is mainly a matter of prompting for the speakers to take turns and to have their lips closed when there is no speech in the audio.
What generally works is to give the official prompt guide (https://github.com/MiniMax-AI/MiniMax-H3/blob/main/.agents/skills/h3-prompt-writing/references/ref-en.txt) to a large language model and ask it to generate a prompt to do what you want.
Description
Removed MelBandRoFormer as podcast do not require audio stemming.
To prevent the lips from moving when there is silence in the audio track is a matter of explicitly stating so in the prompt.