MM H3 reference video ref2va lora
** im working on a much more powerful version, stay tuned
MM H3 lora trained on the reference to video model (ref2va), fun little project for me so I could make stupid videos. I'm using euler/beta. I would be cautious about mixing some other loras (test first) because when I try them they make videos look poor quality or cause smearing. I would also test without turbo first (I dont use them). use only ref2va trained loras. Don't forget with ref2va you can supply short 512x512 reference videos for your scene, that really helps a lot.
FYI: I dont use turbo lora, for various reasons, and I dont test loras with turbo loras enabled. I think it makes the distilled lora training related issue worse, rendering the lora useless.
Grok H3 Prompt Skill: https://grok.com/skill-link/6dc971af0dd0458b12ef523fa62696bf
How does a reference model work?
It doesnt work the way you are used to with traditional models. This is a reference model, it literally means you need to give it reference images because the lora and the model itself is trained on reference images. If you want a nude woman to suddenly run into the scene, give it a reference nude woman. If you want someone to look over and see two people fucking give it a reference image of what fucking looks like. Then let the lora animate it. The lora was trained on "take these reference images" -> "generate a video using them". it was NOT trained on text to video, it cannot generate a cock out of thin air, it needs a reference image to animate it. Otherwise use the fl2va model because it can be trained on pure captions. So if im doing a handjob scene, I might give it the main photo of the scene, a reference close up cock image, a reference oily tits and reference handjob technique photo. Doesnt mean I need all those, I could get away with one image but I cover my bases so the video turns out. The job of this lora (and what it was trained on) is to take images and animate it, it was not trained to create anatomy that doesnt exist in the references. It can sort of do it but it will look incorrect (deformed cock head or something because the angle of the cock is different than what it may have seen).
I put in the comments for each of my v1.2 videos how I created it. I uploaded a "what does a nude woman look like from all angles" photo to the gallery, reference it in your generations so the model doesnt have to invent anatomy as the camera is moving around. https://civarchive.com/images/140155060
you can even give this model video examples of how to fuck, or audio references of what it sounds like, the sky is the limit if you are creative. But its a bit of work compared to the simple fl2va model. this lora was trained on how to make reference images move, hopefully that helps.
You need GOOD prompting, use the grok prompt skill below unless you want to write 1000 word essays.
the two versions of this lora are described below, I would try 'sexytime' variant first.
Join our Discord: https://discord.com/invite/GCrr4Cj3D
WARNING: for the AfterMidnight ref2va lora you need to use euler sampler and beta scheduler, or you will get weird crap like audio artifacts or turbo tits. Listen to some of my videos and you can hear odd audio noise or see strange crap, thats from me using res_multistep. dont be a like me, cheers. this video uses the correct workflow https://civarchive.com/images/140029319 :)
Make sure the lora passes to your Basic Scheduler and Guider nodes or nothing will happen
Workflow: download video and open in comfyUI https://civarchive.com/images/140064227
Lora "Flavors"
there are 2 flavors of this lora (same dataset different training styles), one only one at at time:
sexytime: strength 1.0. Was trained for sex scenes and coherent motion, probably closer to what you want.
softer version: strength1.0, wasn't pushed on motion as much, focused more on detail. I run at strength 1.0. Was trained on detail and surreal style, think like crisp outputs and fantasy (flying dicks?)
Model: MM H3 ref2va bf16/int8
Training: mixed bag of 1500 videos of various sex acts, close ups etc
this is my first attempt at a ref2va mm h3 lora. Still working on my training pipeline for this one (I dont use ai-toolkit)