Please read the relevant version info on the right side.
I think some experimentation is still needed with the generation workflow especially on version 2.0
If you get good results please share your settings.
This model was trained 4000 steps with the provided musubi training files running for about 5 hours on a 5090 and consuming around 25GB vram. I used 56x145 frame clips at 24 seconds.
These I got by running a script to splice longer videos (attached to the files if you want to use it, just put it in the folder with the long form videos and run it with the path to the output dir and you get the cuts).
I've included the training files in the training data, I used this to train: https://github.com/AkaneTendo25/musubi-tuner/tree/ltx-2
(some comments have stated that branch ltx-2-dev has more bug fixes so going to try that next).
I definitely feel it's undercooked but the loss hasn't gone down too much last 1500 steps so I won't push this run any further. Specifically I think training at rank 8 seems to be a bit too little for it to learn the motion and also 1e-4 might be a bit small. Going to try with Prodigy rank 16 for the next iteration but regardless, I absolutely love the quality of the generated image as well as the sound seems to have trained remarkably well.
The workflows for both the starting image and the i2v are my own concoction and are attached to the image / video.
If you have any suggestions on the training / workflows feel free to share.
Hope I didn't miss anything.
Peace!
Description
FAQ
Comments (14)
well.. i would give it a shot.. but every preview you posted has the same woman, not a good sign.
I mean, nice choice none the less.
Every LTX2 preview for T2V has problems, and they're probably cherry picked before even being posted. I haven't tried LTX yet but this is not engaging at all most of the time. Or the people who know how to use it just don't share their knowledge, maybe..
@GlowingGuardianGirl Unfortunately I have to agree... the key most likely is in the actual gen settings. I've noticed massive differences depending on even slight changes and it definitely would be nice for the people who cracked the code to also share a bit of knowledge.
@tensor_fanatic It seems very cool, but at long as there's not a proper and clear tutorial from one of those guys who made it work, even trying sounds like a waste of time. For now...
@GlowingGuardianGirl I dont fully agree with that. There are many people who each know some things. If we can just collaborate we can bridge that gap without waiting for any one person to lead the way.
Its only one view one camera angle
@tensor_fanatic yeah, I kind of agree with this, I've done a lot with ltx2, posted my settings and experiments in a lot of my videos i've posted.
The issue is, if I did NOT have multiple gpu's, I wouldn't touch ltx2, its got to be dog shit slow if you don't have hardware to keep text encoder, model, and vae loaded completely.
If it helps anyone, I am running dpmpp 2m sde huen gpu, with beta11 in comfy ui.
only need's 5 steps. and on a 4090, sampling only takes 10 seconds at 640x480 per 5 seconds, full start to finish 23 seconds with text encode and vae decode.
but again, I'm not swapping models or quantizing, so, your times WILL very.
but, and i can't stress this enough IF YOU HAVE A RTX 4080 OR ABOVE, USE THE DAMN FP4 MODEL, it's only 12gb on card, no speed increase, but vram savings is huge. may even work on a 3090, not sure how the 3xxx series does with fp4, but FP4 DOES WORK ON 4xxx series.
if you want to do talking, you WILL NEED a second stage upscale, but, something I did learn, don't let second stage reprocess audio, use the mask to lock audio after first stage. ltx is constantly adjusting, mapping lips to sound and sound to lips at the same time, if you lock audio for second stage, it takes a lot less time for second stage to lipsync propery, my second stage is only 3 to 5 steps depending on what I am doing.
but, no, ltx2 (from what I can tell) wont lipsync below 1280x704 unless your really close to face, lip sync isn't as fined tuned for lower res.
i do my best to share what I have learned when posting videos, its just not in all videos. but my workflow is in my most recent videos.
@MrReclusive666 Hello there. Would you mind making an article about your method with the attached workflow? See, you're helping a few people here with those tips, but it'll probably stay invisible for most of the community who don't dig in comments sections. That would be pretty neat and helpful. This site would deserve a proper section to share the knowledge... Thank you for the few tips, saving this in a text file if I ever start LTX2! Cheers 🙌
I posted one I generated today with I2V actually tried many others and it works great, at times the fingers move off to the side too much, but normally fixable with a new seed. I also tried some wih some added loras, not bad at all.
Every lora ive tried in T2V completely came out wrong because of how ltx is trained. Wan was easy but for some off reason ltx is so damn hard to get it right. Its like the lora and the mail weight are fighting each other as the images are being generated and never work together correctly.
nice movements and sounds, but she can't aim, she is always fingering her thigh :D :D :D
Yeah, the problem is ltx doesn't know what a pussy is at all, so to him fingering thigh is good enough XD
@tensor_fanatic "vagina" is working better and if a finger is already inserted in the startframe it helps a lot also
