beta 5:
Use TURBO-hybrid_int8 by default. Other options are for experimenting. Full versions here.
beta_5 is built with a different normalization technique. It's also built off 7 different concept-grouped grafts and consensus merges from over 20 Loras, no full direct lora merges. It allows for a non-turbo version which is available. The files with TURBO have a hybrid turbo-delta fusion baked in saving 4.2 gigabytes of memory instead of having to load both ref/fl turbos. Other concept Loras will also load on top more readily, usually needing lower strength 0.2-0.6. For t2v and even some i2v you should 100% be loading mystic_v4, anatomy enhancer, or a related concept lora on top.
Also consider using the non-turbo with PDD, 6 warmup and 4 PDD turbo steps.
Prompt is the entire key to quality and success. Any issues you'd want to blame on a model or workflow, you can go ahead and take a look at the prompt instead. Describe things cleanly and literally:
❌ He puts his penis in her pussy, make hot sex
✅ Live-action pornographic explicit sexual and sensual intimate POV recording: The man slowly moves his lower body forward with his finger on the base of his penis shaft. The tip of the penis slowly disappears into her pussy hole and her wet labia open allowing it entry. He keeps swinging his pelvis forward until his crotch touches hers. Then he backstrokes and starts a repeated thrusting sex motion. The sex act makes her whole body recoil into the couch, bouncing her breasts wildly. She stares lustfully into the camera.
Prompt in temporal sequence on a linear time flow. The longer and more embellished the prompt, the better the output will be usually. 100% use some kind of LLM enhancer; Grok is the one I use since he has a skill preset for reference prompting.
The version just labeled 'hybrid' is non-turbo. Non-turbo full step audio will always be better than audio on turbo versions.
Some working sampling setups:
er_sde/beta57 4-6 steps (seems to be best preservation of style and reduced drift)
res_multistep/simple 6-9 steps (good motion quality, beta will be the sharpest fast motion at 8-9 steps but will give a plastic/burned look)
LCM/simple 6-8 steps (best turbo audio)
Euler/simple 4-8 steps
Next version should be mainly powered by the first round of Sulphur training. From what I've seen it will fix most missing audio and missing concept issues.
Updated - Full credits to these lora makers for having deltas involved in a consensus merge in beta_5 version:
alcaitiff, MisticRain69, diogod, FourBunny, HearmemanAI, tazmannner, simonishere, QualityControl, blo01, ComfyTinker, kermitfrog1202, qdr1en, coachbate, salttaro, alternative_penguin, misterxrex, freek22, definitelynotadog
beta 4:
rebuilt on beta3 config with some loras changed out for newer versions. Turbo was consensus merged to construct a ref/t2va hybrid turbo lora. That's the key element to the merge, there is no non-turbo version. The non-turbo version is bad and doesn't work since the turbo weight is normalizing the merge by it's strength. That custom turbo merge will need further improvement. This one preforms video, reference, and motion well in 6-8 steps without the issues from beta3, but audio needs shift configuration.
Use sampling like Euler/simple 8 steps with sampling shift - 12 video/ 7+ audio. LCM/simple or beta with 6-8 steps and no shift can also be better for audio and drawn styles.
Audio is lackluster and it's becoming somewhat apparent that H3's integrated audio is not good, like terrible actually. Any future multimodal models should avoid integrated audio if they intend to open source. I have almost no control over how the audio works inside the model. Don't post about it. I focused on motion and prompt response and of course when I get those working well, the audio ends up bad, go figure. I'll look at what kind of different turbo configs can enable more audio crispness to come back or likely will have to wait for Sulphur to replace MysticXXX which is contributing to the audio quality drop.
beta 3:
Rebuilt on the delta1024 reference hybrid model. Does t2va and reference. Treat i2v as a single image reference, don't do i2v prompting. Silver's merged turbo is integrated and tuned for 6 steps with simple schedule with the extra _emb layers tacked on to the model, not sure if they're needed.
Samplers like er_sde or multires or other turbo sampling setups work. Best imo is just er_sde/simple 6 steps, no shift, no spectrum, no cache, only a comfy_kitchen backend selection node. 3 video references in a 15+ second outputs can be done in under 10-15 minutes now on larger cards with no extra cache or quality hit needed.
Many (like a lot, all the good ones) on-site loras were fully combined into a consensus weighted merge with ranked drop-out to form an initial part. That merge is put against new, more powerful Wan and LTX grafts as a blend/reshape that uses the loras to consensus shape the grafts, but it also allowed some of the better loras clean pass-through. This is not a linear list of loras just merged. The main element that can present most is probably MysticXXX which was given the most pass-through weight since it's just good--and all 3 release steps of that are inside it. However, they're all-combined with a ton of other loras with agreement and consensus of shape and then it's only reshaping the grafts parts. The results of that are the actual weighted loras that are used to make the model.
Due to drop-out and consensus merge, pretty much all of the loras can all still be used easily on top if needed, and might work better even. All of this was only done to create a large rank dummy lora similar to what sulphur data will look like as a lora or extracted lora so I can start looking at how to apply it cleanly.
It's definitely not a few on-site loras that are linear merged, uncredited, and then renamed with some emojis. I'll only do this until sulphur tuning steps are in my hands and I can work with more targeted and shifted stuff, plus I was tired of waiting and I wanted fast easy i2v.
There is one quirk of the hybrid h3 usage: don't use it i2v. It should either be always used in reference prompting mode, or t2va prompting mode. Even if there is just one single image input it needs to be used as reference and prompted in the ref2va format. If you run an underdeveloped or manually written prompt you will get odd outputs, random camera changes, and blue lighting color shifts when you use the i2v prompt style.
Full credits to these lora makers for being involved somewhat in beta3 version:
alcaitiff, MisticRain69, diogod, FourBunny, HearmemanAI, tazmannner, simonishere, QualityControl, blo01, ComfyTinker, kermitfrog1202
Beta2 and previous:
This started a finetune-by-graft. Or maybe a GST - grafted shift of transformer (cross-architecture). I made both up, because there aren't any projects that have done it that I know, except one reddit post that made me look into it. I experimented with Wan and LTX on the side which led to the initial LTX Eros scripts that became what powered this, all before H3 ever came out. It seems like unified unbiased models like MMH3 can technically take attention influence from any other DiT without breaking if done correctly. Anima, Krea2, LTX, Wan2.2, Flux1 were all tried out, configs tested, about ~40 hours maybe of working in the dark without any paper or technical documents from Minimax. Eventually I developed linear-magnitude blend application and specific block and head gate targets allowing for a smoother graft on an attn-triplet-unfused version of H3 output as a patch file. That sent to lora extraction, then merged to checkpoint at taste. This is a merge but a merge of LoRas I extracted that interact to produce this current shift. I saved 5 ponds of water by recycling data in a few minutes on a single card instead of toasting a server up.
Turbo not recommended yet for i2v, especially when used with other LoRas. T2V use with turbo is better. Use 20-25 steps normal sampling with no dialogue, 25 steps with dialogue along with cache nodes and attn modes. More steps over 25 are not neccessarily better, and can be worse. Use full int8: int8 model, int8 VAE (if it doesn't crash comfy), int8 qwen3vl along with current cache or attn mode nodes. For smaller cards: quants, macOS ports, and Wan2gp support will likely appear on huggingface but not from me.
Known quirks:
Audio difference v.s. Base - This model's audio changes come from attention shifts seeking alternate audio pairing. Attn triplets were unfused before graft, both standard and triplet q_attn was grafted holding about maybe 10-15% audio influence, attn_k was frozen and MLP fc2 layers were untouched resulting in minimal audio interference. This was the main issue with the entire transformer graft and protecting audio. However this version is slightly louder overall than the base model.
Low resolution detail smearing - Some finger digits and fine motion will smear more at low resolution, also a problem in base model. As memory use gets more efficient increase resolution or work on the composition to get around it.
Odd outputs - This can attempt certain concepts more liberally than base model, but that can lead to some undesirable outputs in bad prompting and certain contexts. Data shift comes from completely different transformers and architecture. This shouldn't even work, so it is what it is.
This model is not dedicated to NSFW as that would violate community license agreement. Sure it can do it, just like base. Any NSFW generations are purely the result of advanced reasoning and tokenization resulting from experimental changes. All terms from the H3 community license also still apply to the users of this version. Don't be a dumbass.
H3 usage still requires very intense prompting for maximum effect. Every motion, every interaction, every sound plainly and fully described. Not with slang terms; with proper actionable words that can be tokenized. Refer to the h3 developer prompting guide, hand that .md file to an LLM or Chat agent and have them enhance or refine prompts along the released H3 developer prompt guide styles using the model's tag system. Certain concepts can be made from pure token reasoning. Consult the prompts in my previews to see certain physical descriptions that I use for some things. When using enhancement give the agent feedback about any issues in the generation and get them to describe motions in alternate fashion, or manually edit it yourself adding a negative like "no X, no Y". Still requires prompt refinement and trial/error for best outcomes.
Sulphur Project has 10k banked to attempt actual tuning. Right now training pipelines are sub-optimal. As always Eros is my personal side project, and this beta was also essentially a speed-run of finetuning, figuring out exactly in what configurations and target areas do you get helpful/harmful changes in the model. This is also a proof-of-concept of what and where to target while leaving the reinforcement quality of base unharmed by being additive.
Description
Reworked turbo and mix.
FAQ
Comments (180)
This is insane quality. The image, audio and motion are the best I've seen so far. Including a fine-tuned and well tested turbo for the checkpoint is a much better choice imo. Otherwise people are gonna find other turbo loras anyway, which will always adversely impact the quality in unknown ways.
It's alright. The turbo isn't just there for convenience it's holding the whole thing together like a gate based on it's strength. As for the model I have way higher concept and anatomy quality as a target these are all iterative steps towards that.
Amazing progress in such a short time!! THE BEST MODEL of 2026!! Thank you for your efforts!
No loras needed.
This is a really great model. I am having some issues with phantom animations in beta4, for example: an I2V of a girl squatting down over the ground, looking back at the viewer from behind, she will bounce up and down like she's riding an invisible person LOL. I'm not sure if specific poses seems to activate some sort of innate sexual motion but I haven't found a way around it when it happens no matter what I prompt. I don't seem to have this issue with beta2 (totally missed beta3 release)
Is it only ref2v or t2v too?
Hybrid = all modes.
waiting for a lora version to help mix with other loras and keep ssd usage down (maybe can work with both fl2va and ref2va?)
A lora version would be messy. It has hybrid and turbo deltas that aren't going to extract at full accuracy and those extracts never give the actual effect of using the model. The model already works both modes, this is all of the h3 modes, turbo, and other stuff combined into 19 gigs already.
From your examples seems Beta 4 prompting is back to normal without need to prompt it as Ref2v, am i right? Beta4 work incredibly great, it's also has the best turbo version you could use right now
I still like ref if you need stricter style or transitions on an i2v type output. I2V is mainly only used for realistic style single concepts or shots.
For best results, is this model supposed to be used with r2va or i2va ?
both but it was advised to use ref2v prompting for fl2va for beta3. treat it as a 1 (or 2 for first to last frame) reference image. I don't know if it changed for beta4.
You can use both. I find the result can be better to use r2va even with one image though if it's going to be a scene with transitions. I2V consistently has issues since it's treated as a loose reference and can give freedom to camera and ignore some prompting and ignore some styles if they aren't prompted. But all my previews so far were done in first frame i2v except for the multi-scene and anime one.
does making explicit prompt like "having s3x", "bl0wjob" etc etc works with your model ?
Yes, "she gives him a blowjob" is the only prompt you need. Doesn't need to be fully prompted out like the base model. Doesn't need the numbers either you can prompt it fully explicitly as long as it actually makes sense.
1,音频确实有一些问题。
2,而且我这里会出现人物说话嘴不动的情况。
3,阴部会偏真人化,哪怕我使用了2D动漫风格的参考图。
Try 2D lora. or Change ur base model. like...Dasiwa?
lcm/beta shift v12/a3 very good result
Thank you for reminding about lcm sampler, twice faster than others!
bata3 我用下来的问题是光线比较暗,皮肤比较偏向真实系,大量的痣,雀斑和粗糙的纹理和皮肤瑕疵。喜欢真实系的不错,但是对于喜欢细腻肌肤的用着比较难受。
Something like that might just always be a problem until I have the sulphur run outputs which are trained on 100,000 clips and is extremely generalized and doesn't introduce skews like lora combining does.
beta4 preserve character consistency really well, loving this version, thank you :D
hybrid 25-49模型+lightx2v ref 0.1模型6步是目前画面+音频+动作综合来讲最合适的,这可能是因为ref0.1模型有空间理解力有关,如果10ero未来融入lora,可以关注lora是不是有足够的空间理解力,而不仅仅是收敛速度快。其他的lora都不太行。有的画面还行但声音和动作都较烂。要测试哪个模型还行就用武打动作来测试,武打动作合理那基本上其他的动作都ok。
Everything else is fine, but currently there is a problem that has existed since the model was developed for other videos. Even for anime or game characters, their eyes are usually pure white or other light colors and do not have distinct irises. After a close-up shot or a scene change, this character will definitely have a very distinct black iris. Neither the prompt text nor the reference materials can save the situation...
Yeah the real issue is MysticV4 influence over stuff like that. Concepts that are nearby what it does get over-influenced by it and it pulls everything into realism. I already think I know how to fix that.
@tenstrip There is another rather strange point. Even though I clearly indicate in the prompt that a character is performing different actions with both hands, in the actual final product and the real-time preview, the character tends to combine the actions of both hands at the beginning and the end into one hand, while the other hand may only follow the prompt for a few seconds during the process.
i think the only proplem is that this model cant genarate anime style nsfw motion
i tried but the mtion always is realstic
any tips?
No it can't really do anime. I already found that out and figured out the fix but just put this version out anyways.
@tenstrip too bad.. stil a good model ngl.
On Beta4 camera is moving constantly even when i put "Static Camera" in the prompt
use "The camera holds a static shot"
@to_MacDonalds_and_beyond I did and camera still moves
This is something from the turbo lora I think. The new one I made doesn't seem to do it. It's that plus some over influenced tendencies. This is why I don't like mixing loras and would rather have the sulphur train.
beta 4 :
- reference image not respected on face and hair.
- white eye syndrom 2 out of 3 times
-camera don't stay static when prompted to do it.
Any reference issue is a prompt issue. There is a camera zoom and reposition skew on first frame i2v but it's activated randomly and it also doesn't appear on reference prompts. There is only "the camera remains still" as a static camera prompt, but usually just not prompting any camera makes it static. Changing the seed once or one line in the prompt can also fix it. No need to focus on a single seed/prompt issue and turn it into and outright thing, this one will prompt differently but haven't figured out how. That motion comes from MysticV4 which I'm gonna fix by pre-merging mystic V2-(V5) with a lot of drop-out until it distills it to just the motion with no realism skew or camera and audio interference.
Yeah, I'm getting a bit worse results on likeness when using ref2va. I have a pretty detailed prompt and the camera angle is static.
@tenstrip I did a LOT of testing with all the Mystics and actually V1 is the best for prompt adherence and likeness, and V4 is that much better at "moves" so I've just gone back to V1 for Ref2VA with very good prompting!
I mus say I am surprised, it works the best i found so far. Thank you! Just the one thing you said is no good. I personally dropped the mystic-XXX v4 for the same reason.
V4 does look amazing, but compared to V3 in REF2V, motion is very restricted for me. (Slow character movents and more narrow motion amplitudes, even when referenced with a video sample).
Sound quality is improved, but doesn't much follow the prompt. e.g. loud moans become soft ones. Also, every other run finishes early without a result which happens only with this V4 checkpoint for me.
I'm using DASIWA's v17/v18 workflows, adapted to skip shift for the V3 model.
V3 has amazing motion that follows the prompt and audio does as well. It just has this slightly waxy/shiny look and image and the sound aren't as crisp.
After all, the progress with H3 and your work is amazing!
How many steps are you using and are you using a turbo lora?
Try increasing your steps to 20 if you're only doing 8. Do a short video or two, play with the prompts and weights.
If you're using a turbo lora (1.1 768p bf16 is what I use) then keep the strength LOW. Too high and you can get what you're describe, shiny, wavey, or distorted imagery and the audio will be worse too.
Also, try it without any loras, 20 steps. You're already using the structured prompt style so you're covered there.
In almost all cases the audio issues and visual issues (like animation) are due to fewer steps instead of the 20-25 that is recommended. I think turbo loras do need more work still.
You could also be a psycho and try cranking the turbo lora up but ... It might result in nightmares! :D
AFAIK, the turbo lora is backed into the Checkpoint, so wouldn't 20-25 steps overdo it? As well as adding another turbo lora? The V4-vide is clean and sharp and the sound as well, it just doesn't follow the motion and audio instructions in the prompt very well. Tested different approaches far to long already, I'll probably stick with V3 atm and see what the future has in store for us. :)
Some people find they run a turbo lora at low strength on top to just increase motion.
@DaddyWolfgang Why use a turbo lora if a turbo lora is already baked in in the checkpoint? Unless i'm mistaken.
@ProvenFlawless it can be baked with low strength
For some reason ComfyUI refuses to create lowvram patches for eros checkpoints, which leads to increased vram consumption. Adding turbo lora at 0.01 strength fixes it.
Very cool!
I think a lot of issues are happening because of people using loras at way too high a strength in there generations. Audio issues cna happen but I rarely run into them. The amount of steps matters a lot, and I've never had great results at 4-8 or even 6-8 steps. 20 is the bare minimum in my opinion.
This tends to eliminate prompt adherence issues, the quality of the render, and several audio issues. RV2VA has more audio issues than F2VA, and it's super important to use the structured prompt method. Like, it's basically required.
All the issues aside, there can be no denying how incredible this model is and it's the best open source video model we've ever gotten. Period.
In terms of audio, LTX2.5 is great for adding audio to videos rendered elsewhere and terrible for video. I feel like LTX2.X was developed as an audio model first and visual model second.
I wish there was a way to take video rendered with WAN/MiniMaxH3 and then layer audio over those videos that can create lipsynch that isn't present. It's easy enough to add audio to things like skin on skin contact, liquids, music, and so on. But speech being added that also targets and animates a characters lips and facial anatomy is a no go right now.
So, H3 is the best there is. When it works great it's just insane. And I've gotten a set of work flows that are for the most part working extremely well for me 90% of the time.
Daswai's workflows and COmfyUI nodes are very good to use as well. I've had good success with their stuff. I just wish they'd train their H3 model to do more than what the stock model can. because I know how good they are at training models and making them just stellar. What they did for WAN can't be overstated, they made WAN2.2 so much better than what anyone else out there could do.
All that said, I look forward to what the community does with H3 and what MiniMax has in store for us in the future. Cause this was a surprising model release for a lot of us.
I really expected a heavily censored model that would just end up in the waste bin like Pony 7.0.
And I want to add, @tenstrip :
You too need to be thanked for all you've done for this community. You're at least honest about what you're struggling with and how things could be improved and such. The stuff you've released has been amazing and your absence would leave a crater in the quality of content this site gets uploaded to it.
Have you used the new "Add Guide for Minimax H3" node in ComfyUI?
you can generate audio separately and inject it at frame 0 and then it generates the i2v or r2v clip conditioned on that audio
@DaddyWolfgang I do this all for my own use, my own prompting style, and my own images and output taste. I just share it. The main goal is to be an offline version of Grok like the one on venice and this one is very close to that. But I do see flaws currently with a lack of anime styles and it has obvious style tendencies that I'm working currently to generalize out of it in the next version.
20 steps unusable in practice, any model require 20 steps is just bad.
Dasiwa is garbage useless without 5090.
It seems that eros beta4 doesn't have very good knowledge in spatial movement in ref2va workflow, it generates dwarf characters even with good references of different angles of character with max setup. In a fighting scene any specific camera setup is ignored, while other hybrid models don't have such problem. It also doesn't understand how to render fighting movements right even with help of combat lora v2. Currently this model only serves fl2va right. I think this issue still stems from the merged-in LoRA being an fl2va LoRA, which has relatively weak spatial understanding. I don't think H3's audio system itself is bad, when I use the ref2va LoRA for acceleration, both dialogue and sound effects are still fairly usable. But the fl2va LoRAs circulating in the community basically all have audio issues. So I think the current audio problems still come down to the acceleration LoRA itself not being very capable.
okay this model has 10 times the understanding in nfsw better than any model out there (even wan and ltx best nsfw's)
no i mean it this thing understands so well like to much
its not that good in anime movement (12 fps hand draw) but it is good in anime with smooth motions
i wish we could have that 12 style movement but it will destroy the whole perpose
The animated style was overridden by other weight. In the next version I already created an entire anime component to offset and preserve it.
@tenstrip okay thats good thanks.
Well i recommend to do a non turbo version also. if you can.
This is a hard model to train after all i dont know why. I tried to do a lora train it was messed up.
I will be waiting for the next version
Thanks ! Could you optimize the a-word in the next version? The problem is that whenever I try to insert something into the a-word, it keeps ending up in the p-word.For example, I wanted to generate doggy-style a-word sex, but it just generated doggy-style sex
There has yet to be any kind of dedicated anal lora. Right now this is just limited to what is available. Sulphur will have it though.
Just wanted to share that the t2v quality is a lot better now. BUT... this really needs skin on skin audio work.
The naked body generated by this model is always with huge breasts. I don't like this.
I think that was always on of the main features of MysticXXX, that one is baked in. But overal it is one of the best NSFW for now. Did anyone tried some loras that could affect this model to have different size of breasts?
@HugMeIntoFace I nerfed mystic v4 intro the ground and dropped-out all of it's tendency after this version. I ran all 5 versions of it together into a combined lora and only lean on it partly in the next version.
@tenstrip You doing GOD work, thank you!
@HugMeIntoFace Upload any size separately as a reference
@maxbogatskiy688 Not helping in my situation. I prefer doing FL2VA using just first one image. But what helps was using flat chest lora on 0.35 you will not get E instead it is nice C (even picture looked little better).
I posted a quantized version in GGUF format, I hope the author won't mind. https://civitai.red/models/2902153/h3-eros-max-gguf?modelVersionId=3281707
While this model is good from a speed/quality standpoint, It has the issue of requiring prompting to meaningfully impact the output. Seed # alters very little compared to other models. Perhaps im doing something wrong.
Prompting to get the output isn't an issue. This has been the #1 thing to achieve anything in any video model. Over 90% of people's issues are the prompt.
The prompt is everything my dude.
beta4做不了扇耳光的效果,原模型是很轻易可以做的。而且女角色说话的时候嘴巴有点裂开太大的感觉。
Hey, coming from LTX 2.3, and trying this is strange. H3 does close-ups great, like blowjobs, boobjob etc. And LTX was bad at that.
LTX is great at penetrating sex, but H3 is not !?
Is it something I am missing? :)
I am not saying this model is bad, it's a general question.
H3 - Look at the glans (tip of the penis) as it reacts to her mouth moving past it : Video posted by MadJoker
So nice. And no lora.
That's a quirk of one of the loras+mix in the version mixing up concepts with the model. The whole blowjob motion is very polluted by some kind of food related eating tokenization in the model. Something like that will be fixed eventually with the sulphur data.
Great job ! Which turbo lora used with this CP ? And which wf u purpose ? Ty
I merged my own, but I have a newer one that is better with much better audio already.
@tenstrip :O :O :O :O :O :O :O :O :O :O :O
:O
:O
@tenstrip Ty ! can u send link plz ?
@tenstrip Audio improved would be huge. this model is great just need better audio and it's top tier.
Огромное спасибо за работу! супер модель ! )
The LoRA merged into the beta4 version isn't performing very well. Severe artifacts show up at 6-8 steps, and faces at medium to long distances are quite blurry.
Same problem, artifacts, especially visible on hands and faces even at high resolution and the camera will randomly move at the start of some pictures even by specifying in the prompt to not make it move. What is sad is the camera move is not triggered at 4 steps, but it is at 8 and even worst if trying with 12-16 steps... It's a pity because motion is perfect, far better than any other model.
I wasn't gonna even upload it. I already know it's issues and why I don't use it and I'm already working on the right version of it. Artifacts appearing are your prompt issue. Faces at medium to long distances being blurry is a core h3 issue nothing to do with what I did.
@tenstrip I'm personally comparing to other models, by artifacts those are "jpeg artifacts" not ghost elements coming from nowhere, hands when moving comes blurry and pixelated compared to other models, so it's not a "prompt thing". I say hands but it's the whole image quality which is worst. The model has the best motion by far but guess it affects picture quality like smoothmix for WAN.
Incress the MP and also use 12 steps
And feed it to a gamma opertion with 0.7 ratio
It will get rid of the blur face on distance
For artifecrs either yout prompt is goofy our you video has low time like 5 sec with movemnt at last seconds
Its good model
beta 4 tend to generate dwarf girls. There peopel said about adding turbo lora but even with adding turbo lora prompt adherence not well.
Why the hell didn't I download this earlier?
Any plans to switch fine tuning the FastH3 model version once it gets native ComfyUI support ?
I tested it myself and it's surprisingly WAY faster compared with the normal version.
Link: https://huggingface.co/Kijai/MiniMax-H3-experimental/blob/main/minimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot.safetensors
Link 2: https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree
Is there a specific prompt to get pov anal sex from behind to work? I could get it to work with the base model + loras. From the front it still works.
You are a savior. Thank you so much for sharing this fantastic model.
I'm not clear on the shift settings for beta4. My H3 workflow doesn't have settings for shift, is there a workflow I can look at it?
Where the hell was this model?Where the hell was this model? I was using the original version init8 But in creating content nsfw I was suffering with him and needed Laura for every situation. But now that I tried this model, without any turbo lora, I used 6 steps, with setting er_sde simple It gave results far beyond what I expected. I tried it with regular content first and found that it was better at following the prompt, and the quality, details, and movement were more realistic and smooth, with much less distortion, especially in complex movements And the little details. In slightly less time than the original, generating a 12-second video on the 480p It takes 180 seconds currently, which is a good difference less than before.And.. Yes, I tried it in a computer nsfw It gives excellent results so far. So yeh, The model is really worth trying. Thank you for the developer's efforts and we await the best in the upcoming updates 🙂🔥
I have questions gay's : Regarding the sound in this model, when I create a location nsfw Sometimes the sound does not match correctly with the situation. Is there a audio vae Is it intended for this or does it just need to be set from the prompt?
Are you using Beta 3 or Beta 4?
@Saxlive I am currently using Beta 4
Just signed up for this and it's working exceptionally well so far! But it seems to have a very strong vaginal sex bias. Even during anal sex it keeps trying to draw another anus behind, which the original checkpoint doesn't do. Hope you can fix it in beta 5, thanks!
I tried Eros Beta 3, but I noticed a strong bias in how it interprets anatomy. Even when I prompt for a standard missionary position and explicitly indicate vaginal penetration, without mentioning anal sex at all, it repeatedly interprets the positioning as the penis entering the anus instead of the vagina. I tried several times but eventually gave up.
How do I get it to insert objects into the anus? I've tried a bunch of prompts, but it inserts objects into the vagina instead of the anus.Try to add the lora "ThumbInButt"
lol lmk if that works
The prompt adherence and quality between bf16 and int8 is night and day, They look like two compeletly diffrenet models
that model could run in a nvidia 5070 ti 16 vram?
dang.. now you got me curiuos..
why did you have to say trhat
@sergiogmx I can barely run it on my 5090, it takes twice as much time to generate as int8, but the 1MP outcomes form bf16 look like 2MP outcomes from int8 :))
@evantopsmith Just rent a gpu for an hour and test it out, you dont have to fully commit
@JoeyDiaz How much time for your 5090?
@loope45Gs 2 minutes for 1080p
beta 4 would be perfect if the camera control was better. amazing work.
What do you guys say about this model? I'm stuck on the wan 2.2 so far, but how does this MinMax perform in generation and NSFW?
I have deleted wan2.2nsfw.
whoa.. wan 2.2.. is this 2025?
@hatt2 Yeah. Have you tried it? I think MinMax is better.)
Wan 2.2 was a great model, but it's clearly obsolete now : 5 seconds, no sound. You have to upgrade your worflow!
@KinkReaper Okay. But what about speed, distillation, quality, dynamics? Are crutches better than wan 2.2?
@TeosKuzen Guy, just test yourself and you'll see the difference. If you're not satisfied, go back to Wan 2.2, but you won't!
Wan2.2 is still better trained with concepts. But its only 81 frames for a few frames of good motion by default. You can also just use wan2.2 clips in h3 as references or use real clips and get better realistic motion effects.
@tenstrip The model is quite heavy, as I can see. Does distillation work well and quickly?
Is the model pruned or full size?
I tried a simple missionary prompt, but no matter how I phrase it, the model keeps interpreting the positioning incorrectly and consistently puts the penis into the anus instead of the vagina. I didn't mention anal sex or anal anatomy anywhere in the prompt. I tried several different prompt variations, but I couldn't get it to follow the intended positioning, so I eventually gave up. It seems to have a pretty strong bias in how it interprets anatomy and positioning.
thanks for your efforts !
Hey there. Thanks for the huge efforts you take to provide us. There is no better model for H3 NSFW right now to my knowledge, just as was the case for LTX. And you share it with us for free. <3
I am not sure if this case matters to you personally at all but here is my single point of negative feedback:
Every version was a huge improvement to me. Also V4 works much better than V3. The only downside I noticed from V4 is it often enough does draw a vagina instead of a penis, even though I add reference image(s). This does not happen all the time but sometimes, to me it seems especially when chaining scenes it gets worse. I suspect it is most likely a problem of one of the new version loras merged, becasue v3 did a better job on that regard. So I am not sure how much you can do about it yourself but I wanted to let you know just in case.
Going back to v3 is no option because v4 is so much better generally speaking.
Anyways thanks again for the efforts and sharing with all the community freely. Cant'wait for the next version.
Both those versions experimented with doing full linear merge and using the turbo lora weight to normalize the entire mix. It's a bit too wild, too hard to mix correctly, and is just a clash of raw lora weights overpowering the model and each other. For the next one I've reworked how normalization occurs for H3 and I'm down to amplifying and limiting certain element targets in the model for each lora. The should make the next one more base-aligned so that it doesn't get wild like those ones, but still has the improvements from the mix.
Is beta5 final or will it see revisions before being released to everyone?
This model understands the prompt better than any other model on this site :- )
dude someone please help me get rid of the jibberish its ruining literally EVERY generation. its awful. how do i ix that?
LOL read the beta4 description Bro. He said don't even post about it. Yeah it sucks ass, NGL.
I fixed that and pretty much all the issues since beta3 finally in the new one. The audio result is a low weighted turbo effect. Some people were stacking the plaugekind turbo on top in his workflow and fixing it.
Great model ty.
The only flaw of this model is that it can't genrate futanari. I love your LTX2.3 model because as i'm a 3DX Futanari comics creator, it perfectly genrate futanari in T2V.
Unfortunatly, the H3 version can't. Hope to see a new version with futa included. And this is strange because i think you merged the last Penis lora into it, and the last Penis lora can perfectly genrate futa in T2V mode.
The quality with only 4 steps is unbelievable. Better than all turbo lora I have tested. You're a magician!
I recommend the following accelerators.
https://github.com/Saganaki22/ComfyUI-sol-attn
and
https://github.com/StanLukuvka/ComfyUI-MiniMax-H3-SPEED
I got a 60% speed increase!!
Hi! Does the quality drop, or is it not noticeable? By the way, would you mind sharing a workflow based on what you mentioned? 🤗✨ Please, I'm really bad at adding new nodes, so I usually just use pre-made workflows haha
wf plzzz
The maximum quality I could make on my 5080 is 1 Mp for 10 seconds, then the model freezes. With 0.7 Mp quality, it takes up to 20 seconds.
https://cloud.mail.ru/public/sq74/DCQQGV8ZL
FinalFrame and Melband can be deleted. I use FinalFrame to make a long video by setting the initial frame of the previous video (and it works!) MelBand is a filter that removes music.
Attention! The nodes overlap each other :) Move them to see what's underneath
Amazing @tenstrip! You absolutely nailed this! When is this coming out of beta to version 1?
Also several recommendations to help you:
1) If you can fix sensual sounds that that would be amazing. Big ask I know but it would put you in legendary status.
2) Please add in more Intimacy and Passion :)
3) More realistic sex movements from the hips rather then whole body.
Thank you for your contributions to the open source community and I love the name Eros. If you keep going you're going to have the best NSFW fine tune available.
Tenstrip, can you please clarify what this means:
"Use sampling like Euler/simple 8 steps with sampling shift - 12 video/ 7+ audio. LCM/simple or beta with 6-8 steps and no shift can also be better for audio and drawn styles."
What is sampling shift ?
Thanks Bro !
It's a node in ComfyUI specifically for Minimax h3.
It's called ModelSamplingMinimaxH3.
It has two parameters, Video Shift and Audio Shift.
They basically deal with how the model progressively removes noise from the latent. As far as I can tell, higher shifts force it to spend more time in figuring out what its actually doing while lower shifts make work like it does originally.
0 shift is an error.
ithink your model have the more polish finihs and is really good following the references characters sheets,, there is any settings i could put to improve the audio when the character talk?
As was written, do not bother to pointing out about the audio, Tenstrip mentioned that will not react to audio comments. It is know issue with Beta4, but Beta5 is already worked on, in other comments here author already stated that this issue and other that was in Beta4 (and some with Beta3) should be fixed or improved. IDK when the Beta5 is coming but we are getting closer every day. (You can see at huggingface repository, that Tenstrip is working on the Beta5.)
@HugMeIntoFace i just use a lora h3 pk parasite turbo instenisty 1.6 shift video 12 shift audio 3, and working pretty good the characters stop inventing words that are gibberish , i used in the seedhunter workflow
Excellent model, it takes face and body reference images well, but the breasts are larger than in the reference. If possible, please fix this.
Any chance for a non turbo version in Beta5 or still a nogo
There is a non-turbo of 5 yes.
@tenstrip Perfect - Do you have a patreon or kofi? Gotta send a few bucks your way
@Sassyphrassy https://ko-fi.com/tenstrip
please dont bake in turbo loras without a non turbo option available.
I'm gonna put this in the biggest headline text if I do this method again. The mix DOES NOT WORK WITHOUT TURBO. It is a blurry distorted mess at 25 steps, and is bad even with other turbos that aren't the one that I mixed specifically for it. The beta3 and beta4 are invasive linear merges without normalization that basically destroy the model. The turbo weight is the only thing normalizing it and it is tuned and tested exactly to the strength of that one. Too strong and there's issues, too weak and it doesn't work. I'd have hundred of comments about it not working if I put those out raw. However, the next mix is normalized with a different normalization I reworked for H3 and has a non-turbo version available for beta5, but it's a completely different situation.
@tenstrip hi, cheers for explaining, sorry if i aggrivated u a bit ahah, thanks for all the hard work
Why, when I generate a video using this model, does the video break down into squares after the second second?
This checkpoint has the best motion I've seen so far. Audio is horrible, as you said but it doesn't do weird stuff as much as some other models
I really loved the motion / and face expressions.
Two comments
1. If there are two people they always try to come close or sluttery voice - hahaha no other checkout does that.
2. Used it with https://civitai.red/models/2909076/endless-minimax-h3-with-endless-lipsync and after 2nd or 3rd clip the character starts to move out of frame no matter how TIGHT i make the prompt.
If this can be fixed in next version, oh boy.. that would be game changer
With the integrated loras, this model can do so many things I couldn't get other models to do. Like show penises in hardcore scenes. I do have a question though. How do I add in loras? The model can't quite do cum correctly still, but every time I've tried to add a lora it badly degrades the quality of the video even at low strength.
you are right, man always make a penis grow from hand and cum everywhere. I think this model embadding a cum lora make it intend to combine hand and penis.
please remove cum lora, people cant cum without hand, i have tried every method to make them cum correctly. but they just grow a penis from hand and cum everywhere.
I know it's beta but just sharing:
Its not too good with futa references. And it's not great at knowing where a vagina is unless the reference image is good/explict.
Great work regardless to OP
Stay with v3. in v4 The camera will always move up and to the right, no matter how it's prompted, or the reference image. This often cuts off the most interesting part of the video...
beta4 does seem to go a big crazy with the camera. Also I wish the EROS plastickness could be toned down. Thank you for your hard work though!
That's what I came here to check. Big crazy. Constant sway and wanting everything zoomed all the way in. Thank you, I'm glad to hear it's not just me.
Try turn off turbo lora and sage, that worked for me
Fantastic model, great movements, efficient dynamics, as always, "tenstip" offers invaluable content to the AI community.
Thanks, friend.💖😇💖
This model is brilliant especially for kissing scenes. Kudos.
it doesn't work on 12GB Vram 4070TI
check your workflow. This model is best,even working on my friend RTX 2060 laptop.
get plaguekindmininmax workflow v7 is easy. 10 sec only took 4 minutes
@nailiang002188 I am using Pinokio with a 4070 Ti and 64 GB RAM, but when rendering, it only produces a gray noise image. Since the Pinokio UI does not allow loading the model directly, I attempted a workaround by replacing the original file with the model, which worked with the Minimax H3 INT8/INT4 ConvRot model but is not functioning with this model. Additionally, I observed that rendering a 15-second file takes 17 minutes. I am not proficient with ComfyUI, which is why I prefer using Pinokio for its interface.
beta5..Please. can't wait.I will do everything to get new version.
and your wish came true!
@limbojunkie I know,and I'm waiting for Master @tenstrip 's command~
it is an excellent model. However, I have one question.
If she is wearing clothes in Shot 1 and undresses in Shot 2, her breasts are exposed from Shot 1 onwards. Could you possibly help me with this issue?
thanks for beta 5 im going to try it
also this version can do anime movment?
Yes I've seen it compared to singularity and base reference. Sits in-between the two of them.
@tenstrip also the fp8 version is lower in quality?
and put a warning that 3000 series rtx cards cant run fp8
and if you have discord server i be happy to join it
But I'm also adding the difference between this model and the singularity model as a lora in optional files momentarily. It vastly improves anime effects.
@tenstrip okay thanks
@to_MacDonalds_and_beyond it's not fp8, but it's an w8, civit doesn't even have the right category for it.
Is beta 5 supposed to be run with shift, or without?
beta 4 is awesome and my favourite checkpoint so far, despite some waxy skin I found hard to manage with prompting sometimes.
Base shift, 12/3. Although for steps lower than 6 you will want way more audio shift to bring it up, 3 is too low for low steps. Sampling can be experimented with a lot more than the last 2 versions since it's more robust and more like the base model and the turbo is cleaner.
Also on waxy skin it's just a reality of the current turbo deltas. It won't be super apparent except on t2v without some style prompting or close-up zooms on i2v that zoom on faces or skin.