V3 update: Included smash cut type transitions. "The camera makes a smash cut..."
As much as I love Wan, if it's one thing it's not good at, it's cuts.
This lora is an experiment in greater shot control, by being trained on hard cuts. Most of the clips are also cut on action, fwiw, to allow for a continuous flow.
It's captioned with "wide-angle", "mid-shot" and "close-up". But I think Wan already handles them well.
Prompt format:
[a brief description of your initial shot.]
the camera makes a hard cut to [resulting shot]
[what happens after]
You can string several cuts. If you do, it works best if you also change the type of shot.
Where's the low noise model? There isn't one. It seems to work well with only the high noise model. I'll still train and test a low noise model, to see if it will improve consistency during cuts.
Play with strength if the scene changes too much. Too low strength will make it a dissolve transition (at least testing seems to imply that).
Using lora's (for the target cut) you can get pretty far.
If the cut doesn't 'take':
Try using less other lora's.
If you use a high resolution, lower it.
Try generating a shorter clip.
Description
It seems to work OK with T2V
FAQ
Comments (25)
I'll just put "|" in promt (description before edit | description after edit), and have the same result without lora.
✨
When I use "|", it's a toss-up whether I get a cut, a transition, or a camera pan.
@Tompte It's a toss-up, even when you say "cut to" or "pan to"?
I was not aware of this, thanks. I'm making an updated version, and I'll try this pattern for the captions.
It's because you're not using the base model, are you?
@neph1 heads up, that's incompatible with NAG. If you like negative prompts, you can't use the | token.
@playnproto266 Alright, thanks. I tried it out (even without negatives), and didn't get any good results, even when finetuned. So I'm sticking to my wordy captions, for the time being.
Slightly tangential note, making complete scenes is precisely what the recently released holocine aims to do. It is based on wan2.2 with a couple trained dit models for the shot controls. You might be interested in testing it out if you have the vram.
Demos:
Git:
https://github.com/yihao-meng/HoloCine
Comfy integration:
https://www.patreon.com/posts/holocine-wan-2-2-142248454
Gguf quants (the base models are 57gb each, <Q4 might be reasonable):
https://huggingface.co/QuantStack/HoloCine-GGUF/tree/main/Full
This is an ad.
@axicec An ad for a paper? For a free downloadable huggingface model? for freely available instructions on how to use it? I don't even see a runpod link or anything anywhere.
@axicec negative. got it working in about 3 hours.. but dont use that patreon workflow.. its "ok" but not the greatest. - Ill post a workflow sometime if someone else doesnt.
this looks interesting i basically abandon tv2 never get what i want. But maybe i should try this.
@axicec Obviously, I am paid nothing to promote this open source / open weight model (which afaik seeks to rival similar pay to use services). So it is only an "ad" in the sense that it is a public notice. If you do not care for the passage of relevant information, why frequent the comment section? Go burn some books.
@NexaMuse Please do, personally haven't tested this model since all the demo samples seemed to be limited to 15 seconds and the size was offputting, but if it can perform well on local gpu it might be worth more attention.
@firemanbrakeneck Testing on runpod right now, legit. but. not out of the box working had to do some guesswork. will update.
I tried is just now. Looks very good and cinematic and the scene controls are very detailed using character references. But boii does it take long to generate a video...Even on my 5090
@fenasikerim Oh, how long would that be? 10-20 minutes per 5s?
@firemanbrakeneck Around 25 minutes for 20 seconds with sageattn
@fenasikerim Honestly, that's not bad for that much consistent footage, if it's indeed consistent.
@firemanbrakeneck https://civitai.com/models/2092660?modelVersionId=2367652
HoloCine is okay, but it takes enormous amount of resources and it's always very low contrast. I couldn't get good visual quality out of it.
@koto2091187 Yeah colors seem to be a big issue. I'm gonna wait for the next version.
@Slowmoe Why wait? You have this lora? :D
Looks interesting, thanks for making it.
Fantastic effort, going to test it today!
Details
Files
hard_cut_200_wan_i2v_high.safetensors
Mirrors
hard_cut_200_wan_i2v_high.safetensors
hard_cut_200_wan_i2v_high.safetensors
hard_cut_200_wan_i2v_high.safetensors
hard_cut_200_wan_i2v_high.safetensors
hard_cut_200_wan_i2v_high.safetensors
hard_cut_200_wan_i2v_high.safetensors
hard_cut_200_wan_i2v_high.safetensors
hard_cut_200_wan_i2v_high.safetensors
hard_cut_200_wan_i2v_high.safetensors
Available On (1 platform)
Same model published on other platforms. May have additional downloads or version variants.