PinkCherry MM H3 (fl2va, Beta)
this is the fl2va model, first frame last frame. not the ref2va (reference model)
Join our discord: https://discord.gg/GCrr4Cj3D
Beta 0.6
im using this workflow v1 with some changes https://huggingface.co/Plaguekind/Minimax-H3
save one my videos and open it into comfyUI and use my workflow exactly with an image first frame as a starting point. you can also do frist frame, last frame, or both. disconnect them in subgraph if needed. if you want to do t2v disconnect the first frame input on the subgraph. use that workflow as your initial testbed. I use int8 unpruned on my 5090 with no turbo. install plague kinds node pack
** beta improves t2v and motion, but its not perfect. still needs further training
** I post new snapshots on my huggingface periodically for testing. for those that want latest checkpoint we are testing. on beta-0.6 at the moment
** Aug 09 2026 - Still deep in training this model, next update should be a good one with lots of major fixes. but will not likely be ready until later this week
** Aug 08 2026 - received some great feedback/comments on weak areas in the checkpoint, and ill work on targeting those weaknesses with further training. thanks for all the input and please keep it comming so I know what you are seeing/hearing.
** Aug 8 2026 - both unpruned (32GB) and pruned (20GB) versions of 0.5-alpha are uploaded. Wont be any new versions of this model for at least a week. as im working on the ref2va model now
Loras: any lora works with this base. We have had good luck using pussy synth, riding and coach bates penis lora at about str 0.4 to address the faults in this alpha release base.
I2V/T2V: Your biggest win will be with image to video, not text to video at this stage of training this model.
Model Type: First Frame/Last Frame FL2VA , however it worked fine in the reference workflow for me. Ill train ref model later
Model Training: Full fine tune of fl2va base model
Training Dataset: 5000+ 4k videos of various concepts
Model Training Target: FFN/Attn blocks only
Training System: Custom Training Pipeline
no guardrails or censorship has been bypassed in this model
This is early still in training model I am working on. Treat it as alpha level still. its the first last frame model not the reference model.
Fair Warning: Expect to find dicks on foreheads, exploding breasts and various other issues typical of alpha models. If that's not your thing, wait for beta.
Training Captions used: cock, pussy, vagina, penis, doggystyle, missionary, slut, moaning, orgasm, wet, hairy pussy, spreading, asshole, tight, fucked, riding, cowgirl, anal, reverse cowgirl, pounding, thrusting, hanging tits, fondling, massage, squeezing, dildo, sex toy, penetrate, squats, flexing, bent over, sucking, licking, tongue, blowjob, arousal fluids, jiggling, swaying, sex position, prone, nude, clothing, saggy, breasts, big tits, legs spread, panties pulled to side, cum, semen, creampie, saliva, masturbating, fingers, fingering, eyes closed, perky tits, rubs, lustful, areolas, pink nipples, mouth open, topless, exposed, voyeur, public, bounce, labia, folds, slut, whore, POV, overhead, low angle, high angle, close up, dripping, precum, round ass, amateur, glistening, facing camera, standing sex position, vulva, slutty, tattooed, pounds, from behind, hanging, stroking, jerking, pants down, panties, cock head, cock tip, wet cunt, passionately , rough, gangbang, group fuck, lesbian, threesome, handjob, double blowjob, throbbing, rapidly, rhythmically, wet lips, cleavage, jiggling breasts, cupping, tittyfuck, teasing, erotic, romantic, making out, blindfolded, thick, crotch, deep, steady, aroused, erect, soft flesh, slamming, messy, slapping, pink anus, bathroom
Description
major fixed to t2v and cum
FAQ
Comments (54)
It would actually be much easier for everyone if you provided the exact workflows that produced the results you're showing in the preview. Personally, version 0.6 still isn't working for me and my friends, and I don't know why. Or maybe version 0.6 still doesn't support text2video mode. I created two generations using the same seed, where the camera focuses on a woman's buttocks, genitals, and anus. And using the first alpha version 0.5 and the new version 0.6, the results were identical. I used the res_multistep sampler and the simple planner with 20 generation steps, without Spectrum Apply MiniMax H3 or any turbo lora.
It's important to clarify: maybe I'm an idiot who doesn't understand what I'm doing, and maybe your model works as you intended.
Please make a post on Discord where the video you generated will be attached to the post, and using which we can open the workflow in Comfiyui.
its embedded in the meta data of the video. its first frame plaguekinds workflow at https://huggingface.co/Plaguekind/Minimax-H3
whats your prompt?
come on our discord, we can help
bro goons with his friends, a true man.
Are you going to release some GGUF versions of your model. My Pascal gpu works best GGUF...
Testing the beta on i2v, good prompt adherence and quality but motion is slow even with prompting faster motions. A few videos the speech was slow as well
yeah would be good to see the video (even better if workflows in meta data) so I can troubleshoot it.
Same here, use the new model then the last version, new model is slower.
@SIMJEDI hrmm, looking at my videos in the showcase would you consider them slow? Just curious if this is a difference in workflow or just the model itself. Im using plaguekinds workflow and nodes, but not sure thats related or not.
@sexgod1979 First of all, thanks for your work. Cool to see this moving along. To answer your question, Yes, I would consider the movement in the showcase videos to be slow. Trying the earlier betas this is something I noticed with my gens too. Haven't tried 0.6 yet though
@sexgod1979
In the new showcase yes I consider that slow. I was using the "Faster! Harder! Shake Harder!" lora at .35 with the previous version and now with the current I have to turn it up to .60 to get comparable movement.
thanks all, got a fairly good idea what causes the temporal dynamics you describe (slow sex). going to go test my theory tonight.
USERS: If you have a slow-motion video and speech, then increase the fps. We don’t want a botched model that is hyper speed due to human error. Please and thank you.
In all of honesty, I’m having a hard time choosing this over the regular model with a NSFW LoRA added to it. The generations are faster (to create) and the face is more consistent with I2V.
Regardless, don’t stop experimenting. That’s how you master a model as a creator and we all need to start somewhere. I believe in you and see the improvements each release. Cheers m8.
I cannot go above 24fps in WanGP
yeah I have a good idea though of what causes it, going to test tonight
@NUGGZ1616 I highly reccomend learning ComfyUI. Even the simple workflow templates work fine. It has templates and now they have an app builder which is basically Forge Neo but where you choose the inputs. I realized that Forge Neo and WanGP make generations due to their lack of tensor compatibility. It makes my 5070 Ti Blackwell 1/3 slower than ComfyUI
nvfp4 or your penis will fall tomorrow
haha. under rated comment
The person in the video talks a lot without being prompted. Even after being told not to. What can I do about this?
1) Important to use english text in poromt!
2) You need to use special template for promt for MiniMaxH3. Try to get it from official workflows and in section Audio describe excat sounds what you want.
@mixailckopp978 Thanks. Yeah, I tried those too, even with a template:
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
But despite all that, I still notice this.
you are not sufficiently describing the sound scape. the model has to create audio and if you dont provide it enough info it will make something up. try like "the sound scape is that of a woman moaning, as he has an orgasm. the wet sounds of sex and the man grunting with each thrust".
Hi I am trying this with the workflow you commented, the workflow is great as well but it seems to have scrapy audio, do you also get this issue? I have try different settings but no matter what it always seems to have the audio really bad.
I havent no, im using v1 of the workflow though and the unpruned int8 model. I can test the newer plague kind workflow and see if it has a problem. maybe try a different workflow or see if the base model in that workflow has the same issue. I see some videos people posted other than me that appear to have normal audio. I would drag my workflow (video) into comfyUI and run exactly my workflow. I just tested a woman giving a blowjob with cum in her mouth, was the best blowjob ive ever seen ;)
Thanks for the model!
It's unusable. Even with the 8-step acceleration enabled, the output is still blurry, and the audio is badly corrupted.
someone just posted a video of a women in lingerie and theres no audio or video issues that I see. my videos dont show that either in the showcase. maybe something specific with your workflow? maybe try dragging my video in comfyUI and running that workflow?
@sexgod1979 I'm using the official workflow. Thank you. Let me test it again.
Hello author, I'm running 4-step inference under LightX2V acceleration, but the video comes out blurry and the audio is corrupted. Does this mode support acceleration at all?
@linzhiyong5231860 I havent tried acceleration, 4 steps is pretty fast. I would definitely try it without the lightx2v first, confirm what the output looks like without any loras or turbos. If it looks normal , then try the lightx2v.
nobody does 4 steps. 8 steps is the lowest you should go even with the turbo lora
ot, but have you started working on ltx2.5?
Thanks great work - also seems to work pretty decent with References
Love you are working hard on this! Pinkcherry was the best! Hopefully someday you can come back to LTX2.5 -- really want to see what a Pinkcherry checkpoint is like on that model. I'm still using it to make killer keyframes for H3.
yeah for sure, once I get this minimax h3 fl2va/ref2va models working the way I want them
I could extra a big rank lora from pinkcherry v1.8 LTX2.3 in the meantime if you like, not sure it how it would play with LTX 2.5, but could be worth a short
@sexgod1979 Focus your sexy god magic on H3. It has so much potential.
Oh, well done indeed. RIP LTX
Easy there. This doesn't destroy LTX the way 2.3 destroyed WAN. Not even close. I will say, H3 is absolutely killing it with FLF, which is pretty much only thing that makes ANY of these models useful, imo. But DiT was the real jump... and this is still DiT. 2.3 still shines brightly.
@Ponder_Stibbons LTX 2.3 is very fun, no doubt. they are all fun!
@Ponder_Stibbons LTX2.3 did not destroy WAN2.2 at all. WAN2.2 is far superior for consistent, high quality sex scenes and videos in general. LTX2.3 does very well for adding audio to WAN2.2 videos. And H3 absolutely destroys LTX in every way. It just came out and it's already doing much more than LTX while also being consistent and producing high quality output at a much faster pace. The audio is better, the prompt adherence is better than LTX to a HUGE degree.
And H3 can actually handle liquids. It's new, brand new. Not sure if you've ever used LTX2.3 when it was new, but it was garbage and still is garbage. I stopped using it the moment this dropped and it was the easiest choice ever. I still use WAN2.2 though, and add audio using H3. But Now we're seeing H3 models and LoRAs that are making it so I don't need WAN2.2 anymore.
I'll still keep the models but they'll be just for archival purposes.
In the FIRST week we've seen better stuff than LTX2.3 can do after it has been around for several months.
@DaddyWolfgang Yeah I completely agree with you, truth be told. I just got so pissed after putting in an absurd amount of time building WAN workflows to make up for the shortcomings... VACE modules, S2V modules, 20 stage SVI chains, and then all of a sudden I can crap out 30 seconds of 4k video in a few minutes, WITH perfect sound and sync. But it's super stiff, and the high noise WAN motion is great. But I find myself using it only for motion now, V2V continuation with newer models mostly. I guess it depends on your niche. I don't do the short stuff, I hate it, I find it pointless.
I've been playing with H3 for exactly 24 hours, so I'm not even halfway through my permutathon, so probably premature to say much, I just don't see such a massive difference in how the models work, both being DiT, that was the essence of my sentiment.
But yeah, definitely agree, extremely sensitive to promptage, pretty much the opposite of LTX. Liquids, yep, no comparison. Awesome liquids, extra-liquidy liquefaction. And holy crap I am getting insane FLF. Which is my personal holy grail, as this is all I have been doing for the past few years. I need my frames to match, pixel to pixel, no garbage, and this is as close as I've seen a model get. There is no reason an AI video should ever have a cut, or be finite. I wish it were against the ToS. That would really clean the site up.
I just disagree that LTX is garbage. BF16 model, undistilled, works spectacularly for me. No upscale stage, straight to full HD or 2K usually. I only get junk when I do the stupid 0.5 scale down just to upscale. Garbage in, garbage out. Yeah it's hard to prompt, but when it works it's friggin awesome. So we shall see, obviously I have a lot of comparing to do.
@Ponder_Stibbons Honestly, i can't get any close to what i'm getting from wan2.2 with LTX, only thing i'm missing is duration and sfx. gonna try to see how close i can get using minimax.
@boulbi78 https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context for long videos
@EddieMurfington Thank you, but it is for H3, 10 seconds are more than enough for what i'm doing. if H3 can manage to handle proper nsfw+sfx it will be what i'm gonna use. (gonna leave this on side in case i need it, thanks again)
woah there, slow down, it has a LONG way to go LOL It still has no idea what penetration and consistency in anatomy is XD
@DaddyWolfgang I agree with your WAN22 remarks. I was just doing som stuff yesterday and was like "Man, Wan22 did this 10x better."
will you make any GGUF versions some time?
I didn't need nsfw but what I was trying to do with charecter creation was not working with other models. (Take a full body image nude, thane then standing in the middle of the room and have the camera rotate 360 degrees around them, then extract the framers to make a charecter sheet.) I figured nothing else was working to try the smallest model here. Figured it would be shit and boy was I wrong. This actually listened to my prompt better than any other model. Maybe I got lucky with a good seed, but it let me do what no other one has done. Sad when a porn model does better than the original with non porn stuff. Thank you.
Damn good work! Can't wait to seethe final.
Dear @sexgod1979 .
I am gladly to inform you that I tested your FF2VA on a MiniMax H3 Reference to Video node with positive results.
With 8 steps with minimax_h3_ref2v_turbo_4step_v0.1_comfy_bf16.
On DGX Spark on 0.5MP widescreen it cost 1 min per 1 sec to render. So... 20 seconds is takes 20 minutes overall.
Not to be indelicate, but how does this file (the 19GB INT8 version) compare to the regular fl2va pruned int8 convrot when it comes to the details of human female anatomy?
我认为这是目前的GOAT