Minimax H3 INT8/INT4 Convrot
Just uploaded both FL2VA & REF2VA w4a8_mixed models quantized by Kijai. (Makes my mixed int4 models obsolete) True int4 model size with int8 activations, near int8 quality, same speed as int8. For RAM/VRAM-constrained systems. Update your ComfyUI to the latest version for support of the new w4a8 (Should be included in Stable ComfyUI v0.31.0 https://github.com/Comfy-Org/ComfyUI/pull/15308
FL2VA - first last (frame) to video / audio
REF2VA - ref_images / ref_videos / ref_video_audios / ref_audios: up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips
Both models can generate t2v (Text to video), i2v (Image to video), v2v (Video to Video), a2v (Audio to video), and multiple references (image/video/audio). But were further fine-tuned/trained for higher-quality outputs for the intended use.
Description
True int4 model size with int8 activations, near int8 quality same speed as int8. For ram contsrained systems
FAQ
Comments (56)
For me this type of model dont work at all. Or maybe is very, very slow... i wait near 5 minutes, but still zero movement... GPU tepmerature was low, so its mean dont render anything...
how many Vram or Ram.
Vram 16, Ram 64.
@herkus_baronas631 Try watching this custom node as your models load, they could be stalling while loading. INT8 and the new Asym_A4W8 works good for me on 3060 12gb 32gb ram. Generation time on the first example of the asym quant model in the prompt. https://github.com/kijai/ComfyUI-MemoryVisualization
@tsolful Ok, would try. TY.
@herkus_baronas631 How'd it go, if it is getting stuck while loading and your on a single gpu add --vram-headroom 2 to your startup bat file
@tsolful I haven't tried it, I haven't had time. When I try it, I'll write.
I had the same issue. I also have 16GB VRAM and 32GB RAM, and I had to lower both the resolution preset and duration to get it running.
I'm using the INT8 pruned model with the Speed LoRA at 0.7 strength, 11 steps, the 0.52 MP resolution preset, and 8 seconds. It takes around 5 minutes or more to generate for me.
Without the Speed LoRA, I was able to get away with 20–30 steps at 0.63 MP, but I had to drop the duration to around 5 seconds, and generations were taking 10+ minutes.
So if you're getting zero movement for several minutes, I'd try lowering the resolution and duration first and just let it sit for a while. My GPU temperature is around 60°C, with about 94% VRAM usage and around 70% GPU utilization in Task Manager.
I actually have more trouble trying to run the Ref2V model.
Speed LoRA: https://civitai.red/models/2837571/minimax-h3-turbo-loras?modelVersionId=3202732
@tsolful Thank you for KJ node for vram and ram. By feelings, i think now work faster. But this new type of quantisation dont work. Ram and vram isnt full, but temperature down, gpu on 99 percent and nothing... so for me INT8 pruned is best solution.
@herkus_baronas631 Hey, if you were getting stuck on the MiniMax H3 Director Guide node when going too high res, you could try adding these startup flags to your ComfyUI config/launch command:
--disable-pinned-memory --disable-async-offload --reserve-vram 1
I was having a similar issue where H3 would sit around 38–40% on the Director Guide and make my whole PC extremely laggy. These flags are supposed to make the VRAM/RAM offloading more conservative. Might be worth a try if you're dealing with the same memory issue.
Давнго пора было взломать эту штуку надоело платить им
никогда не платил. Но да, модель, конечно, мощная. Даже ужатая версия. Этакий карманный Seedance.
just as a heads up.. both w4a8 checkpoints you uploaded are actually the ref2va checkpoints
For me the they are correct, and fl2va is the 11.68gb ref2va is the 10.96gb
Edit: They are different files, contacted support as they are downloading as the same name
Is the latest Asymw4a8 model supposed to be slower than your first int4b_q model ?
It doubles the time it takes to generate
Many thanks to the developers for creating such an outstanding model and open‑sourcing it. I am still in the process of learning. During practical operations, I find it difficult to control the fine details of camera movement. Could the developers and experts in the community please take a look at this design? Is it feasible to implement this feature as a plugin? I created a somewhat flawed schematic diagram using nano2. https://civitai.red/images/139169154
All 3 files in the "Required Components" section fail with: {"error":"Unauthorized","message":"The creator of this asset has disabled downloads on this file"}
Yea its on civitai's end, when i upload the files from my rig the files get linked to the official civitai minimax model page which is set to generation only (with no option to unlink the files), so the files are unable to be downloaded as of yet. I'll update the description with the huggingface download links. Forgot about it as someone else reported this issue to me a few days ago, Thanks for reminding me
I tested this but for NSFW its absoluteley NOT WORKING
WAN 2.2 is still superior for NSFW scenes, like s*x, penetrating, cum. but H3 is better on undressing, flirting, seducing and erotic talk.
@RavirKun Since it’s still in the early stages, a lot of LoRAs haven’t been trained properly yet. This model is far more powerful than Wan.
@RavirKun @junweifeng Yeah, in a few weeks, i would expect nsfw version that will leave wan in the dust... plus it can talk and moan!
I picked up all the required components from the links in the ComfyUI template. Just came here to look at the gallery and comments. Works great. The turbo lora added requires less memory by just a smidgin overall and obviously faster but looks cartoony in comparison to without. It's worth doing the 20 steps!
what about sage attention? have you compared with and without? i'm using both (with the turbo lora) and it's fast but yea, i notice the quality
@bigmoneypat363 sage is worth the speed boost! i hear nothing but good things about it. i never noticed any quality loss when i originally got it
@bigmoneypat363 Ok, I had that deactivated because of a version conflict in the python module at some point and I hadn't needed it (spent a while reinstalling comfy and python entirely after it failed and somehow left my installation trashed) but longer H3 videos have been making me want to try turbo again and I just forgot about it. Updated Comfy recently and sageattention works fine now. I think the quality is certainly much better than without the sageattention turned on. Trying turbo lora with turbo sampler and tiled vae decoding because I was getting crashes at the decoding step on longer videos. Looks good so far! Thanks.
@bhopping I am doing some more testing and there is definitely a quality loss from 20 steps without sage attention. I was using 25 or 30 to get a bit better results really. But for now I'm happy with the turbo 4 step with sage attention which is several times faster. It's definitely better than without. I really didn't know that sage attention would make the quality with the 4 step turbo better I though it just provided speedup overall. Thanks!
@bhopping @shantar6 Experimental int8 sol attn, even bigger speed boost, last position and keep sage attention patched, when sol attn isn't active sage is. https://github.com/Comfy-Org/comfy-kitchen/pull/117
I found a fix when getting stuck on the MiniMax H3 Director Guide node. Add this start up flag to your .bat or .ps1 file --disable-pinned-memory. Have fun generating in 2K now!
for me, it is getting stuck at 0% on the sampler :( also, what file are you reffering to? what location in comfyui?
@KiraNugget Edit your comfyui windows portable run_nvidia_gpu.bat file, here's a node to view model loading also https://github.com/kijai/ComfyUI-MemoryVisualization
.\python_embeded\python.exe -s ComfyUI\main.py --windows-standalone-build --disable-pinned-memory
@tsolful will try that, thank you :)
@tsolful it seems to work, I just had my first step after 3 minutes, earlier I ran the same video and after 45 minutes it was only at 75% before I gave up.
@KiraNugget There is a low vram workflow provided under AsymW4A8 model page, i use that workflow for my local 3060 12gb 32gb ram rig. Try lowering the MP to 0.4 5 secs
@tsolful I have a 4070 ti with 16 vram but I don't know why it takes so long recently. I'm using the turbo lora, 1 megapixel. with the pruned fp8 model with sage attention enabled in Dasiwa's workflow. I don't know what i've change but the first days I had 15 sec clips generated in 15 minutes... now I just did a 10 second clip in 20 minutes. 15 sec seems very slow still, I haven't had the patience to finish one as the sampler seems to get stock at 0% still :(
@KiraNugget I have 4070 tisu with 32gb sticks of ram. i am able to push it to 2.10 MP - 2K res preset, 3 seconds duration, 12 steps with turbo lora at 0.7 strength with the ref2v int8 pruned model with 2 reference images that takes 20 min to generate but the quality is nice at least. i am also using sage. what model, steps, duration, and MP are you using so i can try ur settings
@bhopping fl2v/ref2v fp8 model, turbo at .75 with 8 steps, 1 megapixel, 10 seconds, sage attention enebled, no memcache, shift video at 10 and audio at 5 like the base workflow and the HMNSFW aio 1.0 lora, also have upscale with model enabled with animesharpv4_fast_RCAN_PU. I tried first to last frame likee that and took 20 minutes.
@bhopping I also input pretty high res 4k images, I don't know if it influences anyrthing
@KiraNugget the 4k image should be fine as long as you have ref_img_size set to "match."
I think i am getting same speed as you on those fp8 models, 19 minutes. i get 13 min, without the upscaler.
@bhopping ok so it seems the video reference I was trying to use earlier was to big then, but i did resize it to a 720p video... guess i'll try lowering the resolution further in the future.
@KiraNugget You can replace EasyCache with Spectrum. You can use the Spectrum Apply MiniMax H3 node instead and get a decent speed boost without the significant quality loss that EasyCache can cause. I didn’t use either one in this test, though the Spectrum node is currently broken for me after I updated ComfyUI.
@bhopping Yeah I have the node, I tried it early on at release and decided to go with the base workflow for now. I will try again raight now see if anything inproves. Thanks for your help!
@KiraNugget np :)
@KiraNugget there has been an update for the Spectrum and ppl said its been improved. theres new default settings for it when you re-add the node to the workflow. there is a slightly smaller vae you can test that might boost your speed called minimax_h3_video_vae_int8_convrot and it works with fp8 model. https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main
@bhopping yeah im using it right now. it works great!
This is magic, ref2va_pruned_int8 (~20GB) on 4070 TI SUPER 16 GB - 8 sec video - 1:48 min... with lightxv2 v1.0, 6 steps. (800x576)
Would you mind sharing your workflow?
@sdsaafdsfdsfdsf I used slightly customised workflow from this post: https://civitai.red/models/2837571/minimax-h3-turbo-loras ; look at "Optional files" section.
Someone made hybrid models combining FL2VA and REF2VA. I haven't tested them yet, but they could be a nice middle ground between the two, using some of FL2VA's quality while retaining REF2VA's reference capabilities. They're on Hugging Face under Minimax-H3-fl2va-ref2va-hybrid-models, with different versions. b30-49 is closer to FL2VA, while b15-49 is closer to REF2VA, with b25-49 and b20-49 in between. Might be worth testing to see if one gives you the best of both.
https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/tree/main
Ooo, interesting find I'll run some tests also
They work good. :)
Hey there, great model!
But i have an issue with the lip sync. it seems the actual lip movement always starts 1-2 seconds too late. Any help would be appreciated
Might happen, the "solution" is to cut the first frames, but 1-2 seconds sounds too long to me. You should state, in the text prompt: "The singer is moving the lips" or something like that too.
Trying to use this model via the included comfyui workflow doesn't tell you that you need a file from Comfy-Org/MiniMax-H3 to get it to work. If you don't have qwen3vl_32b_minimax_h3_int8_convrot.safetensors then you'll need it before the workflow works!
Why didn't you include a node for Turbo loras?
Can someone help?Every time I make a i2v nsfw film like side sex position,the character always change the face.And lots of time the face become fat.
I lose the consistency in face.
What's wrong with this situation,is the turbo lora make this happened?or I had the incorrect prompt?
I make 6s output(for 6 steps)
with Fl2va int8 model