V3 Update
Now actually the fastest HQ audio + video generation workflow. See for yourself in the examples. I have intentionally created shots that show distant faces, because that is still a Minimax issue and no workflow issue currently. I can currently create these 8s, 1 megapixel videos in 3.5 minutes which is INSANE if you ask me, while keeping good audio!
I'm still experimenting with different scenarios, tweaking stuff like first pass generation (0.2 - 0.4 MP), aswell as step count for the first pass
WHY 2 MODELS AND REFINEMENT AT ALL??
Yes, valid question. My answer: creating high resolution videos with the Fast H3 model takes much longer than this workflow, and the quality of the short Taomate turbo workflow is outright bad with low steps. The Solution: We use the power of the audio quality and base movement of the Fast H3 model while then using the ultra-rapid speed of the Taomate 3 step turbo lora to refine and enhance the video to High Quality!
PLEASE let me know how this works for you!
What changed: Completely reworked the 2nd pass to use the base T2V model from H3, you can of course try and use other ones. It is important to use the Fast H3 model for the first pass.
Base Information
This is my very first attempt at making a FAST workflow that runs on my 16gb + 32gb System RAM Setup and I'm looking forward to see if you guys have the same results like me.
This workflow uses 2 passes with 2 different models - the Fast H3 model + another normal T2V/I2V model. The first generates a smaller resolution video + the final audio, using the Fast H3 base model. The second pass refines the video into a higher resolution using the Taomate 3-Step Lora. Also Attention nodes are in use as well as Previews for the first and second pass seperately, so you can already decide if you want to keep generating the higher resolution video or not. This combination ensures good movement aswell as the good audio from the first pass while enabling high resolution generations in only 3 refinement steps at 0.35 denoise.
Please try it out and let me know what you think!
INSTRUCTIONS:
Use the green nodes to configure your resolution, aspect ratio & duration
Keep the aspect ratio the same for the first pass and second pass resolution selector
Add your desired LoRAs in the purple node
Adjust to your desired VAE's, text encoders that you might already have (I'm using a new KJ video vae that is smaller)
Don't forget to adjust the filename and path in the Save Video node to your prefered syntax and location
MODELS
Custom Nodes Links
GENERATION SPEED:
On 5070Ti with 32gb System RAM:
8 seconds - 1 MP resolution: ~3:30 minutes
10 seconds - 1 MP resolution: ~4:30 minutes
10 seconds - 0.7 MP resolution: ~3:10 minutes
Description
Made some small adjustments to make better quality.
FAQ
Comments (10)
Please share your experience/results with me, so I know if I should keep sharing these or not.
I'm 4060ti 16g + 32 g 2400 MT. I'm downloading and give it to try
@lilifirenotalone542 looking forward to hear your experience with it!
I would love to hear how this runs on systems from you guys with more or less RAM available, aswell!
Please let me know of any issues you encounter with the V3 version!
I'm really looking for some feedback here :D
Thanks this looks promising. I will try and report back.
Good workflow, much better than the 0.5 one-pass. But I cannot find the SolAttnMiniMax node. and ComfyUI Minimax H3 Latent Upscaler is Comfyui_Minimax_h3_latent_Upscaler-Plus
@vrykol Thank you! Trythis for the Sol-Attn node: