I didn't bother pixel-perfecting the nodes or lining them up neatly—as long as it’s clear what goes where and why, the rest doesn't really matter. Wasn't aiming for aesthetic perfection, and honestly, didn't have the time either! =)). Recommended page file size: at least 128 GB. Generating 5 seconds of 1920 by 1088 at 48 fps on 5060ti 16 + 64 ram takes about 500 seconds, 490-530 seconds of 1920 by 1088 at 24 fps on 5060ti 16 + 64 ram takes about 290-320 seconds =)) For optimal performance, NVMe speeds of 3000 MB/s to 3500 MB/s + . The best choice in terms of speed and quality | 0.5 | 16:9 | 1920 x 1088 | 9 | 3 |, perhaps it is still possible to optimize the process and finalize it, but there is no free time yet, it is even better to wait for the optimizations of the model and the associated nodes =))
v3.0 I2V MiniMax H3 + LTX 2.5
The upscaler has been ported to LTX 2.5
Certain settings have been changed and optimizations added to accelerate the overall process
Toggles are divided into groups, and options to quickly change step settings for upscale have been added
for optimal processing, tailored for smooth video or with a higher tendency toward dynamics
Settings for working with Lora for 4 steps have been factored in, with descriptions added for this specific Lora
An option to crop and fit images directly inside the workflow has been added "for the lazy", with the new Crop (OreX) cropping functionality, aspect ratio tests were conducted in various options, such as 34:9 and 9:34, and different combinations rounded up to 32 pixels; in all options, the result was good.
An example of an alternative model that works well in Ref2VA mode is provided
as well as a toggle between I2VA and Ref2VA.
Ref2VA uses a lightweight model that also performs quite well in I2VA
it was tested with up to 3 references (image + audio + video) and the model handled it quite well.
Description
Initial release of the MiniMax H3 + LTX 2.3 pipeline. Fully tested and ready to use!
FAQ
Comments (7)
The SSD killer XD
Just a little bit, provided that there is little RAM, by the way, the swap file can be used on any device, even on HDD =)), the speed is needed to reload models with multiple generators without overwriting ssd cells, which does not affect its wear while maintaining an adequate temperature =))
I will try it just for the giggles. also you sure like your women wide like the barn doors, so you never miss when you try to get in...
It's all humor, body positivity =))
First of all, thank you for sharing this workflow — it works wonderfully in T2V and delivers excellent results. The only issue I’ve noticed is with I2V: the face tends to change quite a lot during generation.
You mentioned earlier the possibility of injecting a fix into the workflow. Could you explain how that could be implemented, or suggest a way to stabilize the face in I2V while keeping the rest of the workflow intact?
Thanks again for your great work, and I look forward to your guidance.
Thank you, there are several nuances for holding the face in the megapixel setting. It is highly recommended not to go below 0.3 if the face is close to the camera and 0.4 - 0.5 if it is quite far away. Also, using the Euler_cfg_pp sampler in LTX upScaler helps to hold the face reference much better, but the requirements at the upscaling stage increase significantly. ... The best option requires the implementation of the 🧑 LTX Face Identity Reinforcer node from the author of TenStrip 10S-Comfy-nodes. I cannot test it at the moment because it requires a lot of time to optimize the process well ... but I plan to do this in the future. For now, I can only recommend what was written above and you also need to follow the 32 rules, preferably cropping the image in advance to the correct aspect ratio.
Sorry, I forgot, for I2V it is still worth using preferably the basic model of the ltx-2.3-22b-distilled-1.1-int8-ConvRot.safetensors upgrade guild. NSFW mode is more prone to changing their appearance, as well as the sigma value can be reduced to 0.385, 0.250, 0.125, 0.0. These are not exact values, but as a test variation, it is also desirable to describe their appearance more precisely, although this has little effect on the upscale stage.
