Hi Gooners. I'd appreciate your feedback and also if you could share your creations on this page. I recognise this workflow is getting LoRA heavy, so thinking of ways to better manage that in future - suggestions welcome.
Version 1.2:
Known issue - 'graph' NRE errors may be seen when attempting to load the workflow. If this happens please downgrade your videohelpersuite to version 1.7.3
Steps adjusted to 2/2/6
Some complex prompts were still causing ghosting on the latest x2v loras, needing another high noise step to compensate. As a bonus, motion is now more consistently fast - i've actually had to start prompting for slower movement sometimes.
Loop mode
A simple way to quickly continue where the last video ends. Can be left disabled (as it is by default).
Color match node to help stabilize the image colors through several video generations
New LoRAs
Gyrating Hips
Fingering Pussy
They do exactly what they say, and very well.
________________________________________________________________________________
The below information may be out of date - will fix soon!
Version 1.1: Lightx2v revision was released today, allowing us to drop total number of steps from 12 to 9 with better motion in the results. We now do 2H*/1H/6L with the first two steps continuing to be run in native WAN at high CFG.
Ultimate Deepthroat has also been added to the nsfw LoRA stack.
The workflow has all of the links you need for the new files.
Consistent and reliable lewd and pornographic videos in WAN2.2 i2v
This workflow leverages the powerful WAN 2.2 model to generate 1280x720p videos consistently without common issues such as slow motion, light artifacts, color blooming and character teleporting.
- Hand tuned step config (2H/4H*/6L*) across three samplers for reliable motion
- Lightning LoRAs reduce total steps needed from native running on all but the first two steps.
- Built in interpolation step smoothes motion with high quality generated frames
LoRAs:
All LoRAs used are linked both on this page as resources but with direct links in the workflow notes.
The LoRAs strengths have been tuned to allow for a very strong enhancement to body movement while still appearing quite realistic.
Performance:
This workflow may have a longer render time compared to low-res alternatives, but the results speak for themselves. The 1280x720 output is crisp, detailed, and mostly free from common resolution scaling issues seen in other methods.
Tips:
ALWAYS start your prompt with a brief description of what the character is doing/can do easily, then your prompt. For e.g.
"the woman quickly glances at the viewer. she jumps up and down excitedly"
1024 pixel length videos take 3m03s with a 5090.
Description
Added looping function, including a color matching node
Added several LoRAs for specific poses and movements.
Updated stepcount
FAQ
Comments (23)
comfyui killed when loading Wan21. i am on a 5090, so not sure if i'm hitting OOM issues. at 20gb/32gb allocated when i get that error
Strange, you could try bypassing the Sage attn nodes in case it's related to that. Have you had any success with other workflows before?
@Diecron312 My issue was OOM due to my WSL swapfile not being set/created. Not an issue with your workflow :). Got the frames to render but failed on the interpolation due to not having CUDA_HOME env variable set. Should have everything sorted now
Have you found a solution to degrading colours if you extend video by reusing last frame each time. The saturation keeps going up by the end of the video, so if you keep extending you need to keep color correcting somehow. I read that adding sigmas could fix this, but the sampler nodes you used didnt allow to add samplers so I tried with another node and tried to plug it into your workflow, but i must have done something wrong, the node didnt have all the options so it could be that. I read it is also caused partially by the lightx2v loras, but without those it takes way too long.
last-frame continuations are fairly flawed overall, but this workflow does implement a color matching node with a reference from the upload photo node which I found helps a lot. It's not perfect though.
Is it possible to modify this workflow so it runs on 3080ti with 12GB?
Yes. The main thing that will help is finding quantized versions of the WAN model and load those instead of the fp8 scaled versions.
Does it work with 4060 gtx?
tried running yoru workflow and this shows up
Loading aborted due to error reloading workflow data
TypeError: Cannot read properties of undefined (reading 'graph')
So I'm unable to get it to work - despite this workflow being the ideal long video work flow because it appears much more simple than other bloated workflows.
Please read the workflow description where this error and it's solution is mentioned.
@Diecron312 Sorry I'm legally blind, and I see it now. The other issue is that the whole nodes section seems like it needs to be connected. Everything is in a giant pile and you have to dig it all out and connect it. I can dot hat - but at this point the workflow just seems inferior to other flows.
@midnightxwaltz914 when the video helper suite is downgraded, close the WF and load it again from file. It should load properly now.
Was given the runaround with sageaware after dealing with the graph issue (thank you for that), TLDR you need python 3.11 and correct torch versions etc, the new comfy pre-packs python3.12 so delete the .venv and recreate it with 3.11 manually, then you can get everything you need. You will need to know your cuda versions, luckily it's easy to get from cmd line.
Yeah the requirements, transformers and attentions are really not handled well. I have managed to get this all running through a docker container so it's all running in a Linux environment now which makes things far easier to manage.
The downside is the looping implementation in this workflow doesn't work in Linux so I need to revisit that at some point
Oh... Is that why I've been struggling to install Sageattention for days? I've had to just delete it from the workflow. The rest works nicely, though no matter how much I prompt for slow action, it keeps giving me extremely fast motion.
@Zerg96 I forget if the chinese prompt is still in the negative, remove it if it is. Translated it refers to still images and things of the nature, which does help generate faster movement and camera movement. You can remove it or update it, even specifying fast movement if you still have poor results.
@Diecron312 Yup, I noticed and erased it, tried specifying fast movements in the negative and slow in the positive but it's still extremely fast for me. I'll fiddle with it some more and see if I manage something. I also get an error when attempting to interpolate so I'll try to fix that as well.
@Zerg96 That actually makes sense, if the interpolation is disabled you need to set FPS to 16 in the video save node. Leaving it at 32 will mean the video plays back at double speed (interpolation would add frames in between the ones WAN generates, hence we can play back more in the same time period)
@Zerg96 And I would definately recommend fixing the interpolation mode (disable torch_compile for faster initial load, but slower generation of the interpolated frames) and leaving it on for the reverse (large first-time spin up time - may look stalled - faster generation of frames).
It's on by default because assuming you do 2 or more runs you should overall save time.
The interpolated output looks far better than native 16 fps from the model, but it has downsides too (ghosting, innaccurate projections etc).
@Diecron312 Thanks! Took me a while but I managed to fix it. Eventually I'll see if I can get Sageattention to work properly too. In the meantime, it did seem to fix the issues I had with movement being so quick (that, and prompting a bit more agressively too). One thing I was wondering, I noticed that even when I specifically prompt for the video generated to end on a specific position, it tends to make the character adopt a similar position to its starting one after doing the one I asked for and I'm not sure how to circumvent that.
@Zerg96 This sounds like the side effect when you ask the model to generate more frames than it knows how. I usually stick to 81 frame gens, upto 101 can be OK, at 121 you'll start to see what you describe - characters 'reverting' to the initial post. Some people theorize that its because WAN is trained on 5 second clips that are reversed to extend each input to 10 seconds - so when you ask it to do more than 5 seconds, it suddenly starts 'rewinding' actions.
@Diecron312 Oh wow, that's super interesting. And you're absolutely correct, I should reverse it to the original 81 frames then. Thanks again, very helpful advice :) I have a last question if you'd be so kind. I'm not sure I understand what the 2/2/6 and 2/1/4 correspond to in the notes about resolution. Because of hardware limitations, I've had to stick to 480x640 since higher res tends to crash my ComfyUI. I wouldn't mind trading in time for higher resolution, alas, I don't understand enough yet.
@Zerg96 the 2/2/6 etc is simply how many steps are taken at each phase. High no lightx2v, high x2v, low x2v:
Normally, wan 2.2 uses two phase (high noise for overall movement, then low noise to get final images)
But the lightx2v loras that allow you to run low cfg (and thus double speed) had a severe side effect of poor movement.
the first 2 steps (high noise, no lightx2v) are 'regular' wan - we're not asking it to speed up for us, but we only want it to create the overall movement
second 2 steps - comes in and refines the high noise further with 2 quick steps lightx2v
third 6 steps of low to correctly denoise each frame
It's basically a hacky way to get faster generations to behave better.
Unfortunately if you are limited on resolution that likely means vram, and your only option is more quantized versions of WAN. I was really hoping we would have nunchaku for WAN 2.2 by now.