Minimax H3 INT8/INT4 Convrot
FL2VA - first last (frame) to video / audio
REF2VA - ref_images / ref_videos / ref_video_audios / ref_audios: up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips
Both models can generate t2v (Text to video), i2v (Image to video), v2v (Video to Video), a2v (Audio to video), and multiple references (image/video/audio). But were further fine-tuned/trained for higher-quality outputs for the intended use.
14.81 GB INT4BQ Balanced-Quality leaning: int8_ratio 0.41 by count, 0.47 by params; int8mm_ratio 0.58
17.27 GB INT4Q Quality: int8_ratio 0.73 by count, 0.75 by params; int8mm_ratio 0.27
Description
FAQ
Comments (103)
waiting for the nsfw one
It's uncensored, but it takes a while to make the videos, although it works well with low VRAM on lightweight models; we need a Lora Turbo.
And so it begins lol.
Can this do nsfw?
Yes for I2V (IDK about T2V) it is quite uncensored out of the box
@tsolful what? u sure
@sadsshit Yea look at the two in the gallery
@tsolful its definitely has censor in it, i tried img to video, it make pussy looks weird, idk how in the galery the make it looks good
@Seii1 yeah u need to be descreptive of the thing he does have to guess when he guess something he does it shitty but if tell him this is like this and this this way it make a wonderful work ask gemini for help and tell him what u doing he tought me it (i still too lazy to do it right but)
Did some testing, it can do breast slapping, it does basic jiggle (breast and ass), have not tried penetrations yet, but so far it understands the very basics, but it failed at breast sucking.
Yes, it's everything you say here I have been using since yesterday
This model is very good, it recognizes sexual organs
私のところではあまり良くないですね…画像に映っていればうまくいくのかな。
Yes! This model—with an 8-step LoRA, realism-enhancing LoRAs, and NSFW LoRAs—is going to be insanely amazing!!
Just from the few i2v gens I have done, it seems very capable.
any chance to get this on 16GB VRAM?
Works on rtx 3060 12gb +32gb ram
@tsolful nice. whats the performance?
@the_remorra959 for 5seconds 480p 15minutes on my rtx 3060 12gb +32gb ram, 5 minutes 1 megapixel 7seconds runpod 5090 128gb ram
Works fine even on my rtx 3060 12gb +16gb RAM. You only need a fast NVMe SSD and to manually set a big (150-200GB) virtual memory file (pagefile.sys). My test here
@tsolful can u share the worflow and teh GGUF's what u have used in it
@Gamert45 I currently use the INT8 model with the default workflow template in comfyui https://comfy.org/workflows/a781503cf508-a781503cf508/, Just uploaded int4 mixed versions for lower memory size at the cost of slight quality degradation
@bnzarev821 can u share the workflow? we have same rig
YES YESSSS
Can 4090 + 48RAM run this model?Thanks
100% the model is working on rtx 3060 12gb +32gb ram although some of the weights are being streamed from ssd
@tsolful Thanks a lot! By the way, how much is the generation time/ per 1 output?
@sekaiwlc07860 for 5seconds 480p 15minutes on my rig, 5 minutes 1 megapixel 7seconds runpod 5090 128gb ram
as a matter of fact im running it on a rtx 3060 12gb 32 ram
@tsolful 5060ti16g+32g有压力吗
@lijia_tu Yes some of the model is being streamed from SSD. Currently working on INT4Mixed models, which will lower the size of the model but also decrease the quality a bit
NSFWではLoRAが必要ですよね?
性器についての理解はWan同様に少し乏しいと思います。
Yeah, genitalia is not as good without a nsfw lora if it is not visible in the Image in I2V
ありがとう!同じ見解でよかったよ。これから有志で多くの人がLoRAを開発してくれると思います!待ちましょう。
@Japamelon 作ったらお互い配布しましょう!!
@tsolful do you know if a vj and pen lora can be made with just images?
LTX2.3 has some serious competition...
A competitor? The Minimax H3 simply destroys the LTX 2.3 in every aspect (general anatomy, complex movements, visual consistency, physics, sound), plus it's omni... Once we get 8-stage LoRAs and others like enhanced realism LoRAs and NSFW LoRAs, it's going to be absolutely fantastic!
More than that, I'd say... I was never fully sold on LTX2.3, but now with H3, I might forget about it entirely. I'll just be waiting for community support and refinements and then make a decision.
Does this mean I should propably say goodbye to Wan 2.2? Or how is it? (I know Wan does not have sound)
@HugMeIntoFace Not yet, for sure.... give this a month or so so the community can do merges and loras, and then re-evluate.. I love Wan2.2, but yeah, the lack of sound has been its main drawback this entire time.
@HugMeIntoFace Absolutely! You can go ahead and delete Wan and LTX right now! Just look at what Minimax H3 does without any tweaking (no 8-step LoRAs, no concept/NSFW LoRAs, and no custom workflows)—the results are already WAY BETTER than LTX and Wan combined! Just imagine what it’ll be like with LoRAs and custom workflows... It’s going to be FANTASTIC! :D
@HugMeIntoFace Go to your LTX and WAN folders and just fucking delete. You won't go back.
@ArttTaku A month!? If you can't tell within the first video you make... LOL
@AI_2_addicted until now it is just slower, and I need around 50GB Ram to render some vids at 20 steps... lets hope the best. :)
@qingyantrong4932325555 @AI_2_addicted Guys I am too old to do impulsive f***s that I could pity later. But I love your enthusiastic aproach. I will look forward what MaxH3 could deriver, in the mean time I will keep the Wan 2.2 warm, it would be heartbreaking moment, if I lost it. I hope that this NEW stuff can continue what I already learn. (I have high demand for consistency for characters and I almost perfected it in Wan 2.2, anything else would be a downgrade)
@tsolful So what is exactly different between FL2VA & REF2VA and one available via comfy download (REF2VA)?
What finetuned means for these download models.
And what exactly these are, am I correctly assuming that these are:
FL2VA - first last (frame) to video / audio ?
REF2VA - reference video / audio to video / audio ?
FL2VA - first last (frame) to video / audio
REF2VA - ref_images / ref_videos / ref_video_audios / ref_audios: up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips
Both models are able to generate t2v, i2v, v2v, a2v and multiple reference (image/video/audio). But were Finetuned/Trained further for higher quality outputs on the intended Use.
@tsolful What does: "But were Finetuned/Trained further for higher quality outputs on the intended Use" mean?
File SHA tells me that files that You uploaded are EXACTLY the same as ones in comfyUi.
So You DL files from comfy and shared them? Is that correct?
@N0n4m3 My bad i was confused I thought you were asking the differences on the models, Yes the int8_pruned models were gathered from comfys huggingface as there was no point converting them myself, Just converted and uploaded int4 mixed pruned versions also
@tsolful put all of the acronyms in the description please. FL2VA REF2VA you explained here.. but R2VA? That's different than REF2VA? What's VR2VA? Virtual Reality to Video & Audio?
If I only have 1 image, can I use FL2VA to generate?
@sekaiwlc07860 Yes
Done some tests and... Wan can 34t $hiet and LTX is nowhere near this one.
Even without LORAs it do produce nice results. Gens are hit and miss but that's expected.
Great model and now my fav.
RIP LTX
HELP!!!Can someone solve this problem?
Node ID: 127 - Node Type: SamplerCustomAdvanced - Exception Type: RuntimeError - Exception Message: RuntimeError: The size of tensor a (75480) must match the size of tensor b (1824768) at non-singleton dimension 2
RuntimeError: The size of tensor a (75480) must match the size of tensor b (1824768) at non-singleton dimension 2
[INFO] Prompt executed in 81.55 seconds
[MultiGPU_Memory_Monitor] CPU usage (97.8%) exceeds threshold (85.0%)
[MultiGPU_Memory_Management] Triggering PromptExecutor cache reset. Reason: cpu_threshold_exceeded
Latent / Image Dimensions Mismatch : Minimax (and similar video models) requires image sizes/resolutions and frame counts to be exact multiples of specific chunk sizes 32 for minimax resolution, 24*(Number of seconds you want)+1
Conditioning / Text Encoder Mismatch: Make sure on load clip minimax is selected
Latent Input Connected to Wrong Node: Ensure the latent_image input on SamplerCustomAdvanced is receiving a valid latent
Which workflow are you using, id reccommend trying the default comfyui template in templates or here https://comfy.org/workflows/comfyui/ should be at the top
@tsolful I use the minimaxH3INT8_fl2vaINT8Pruned model, and only 1 input image. Am I use the wrong model?
@sekaiwlc07860 No it is intended for I2V with optional last frame input, use the image to video workflow in the comfyui templates
If I choose the model, and only have 1 image connect first frame(no last_frame image),can it work? or I need download the REF2VA MODEL
I use the comfyui template workflow,But already the same error:
[ERROR] !!! Exception during processing !!! The size of tensor a (59940) must match the size of tensor b (1451808) at non-singleton dimension 2
[ERROR] Traceback (most recent call last):
File "C:\ComfyUI_windows_portable\ComfyUI\execution.py", line 545, in execute
output_data, output_ui, has_subgraph, has_pending_tasks = await get_output_data(prompt_id, unique_id, obj, input_data_all, execution_block_cb=execution_block_cb, pre_execute_cb=pre_execute_cb, v3_data=v3_data)
File "C:\ComfyUI_windows_portable\ComfyUI\comfy\model_sampling.py", line 97, in noise_scaling
return sigma (s noise) + (1.0 - sigma) * latent_image
~~~~~~~~~~~~~~~~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
RuntimeError: The size of tensor a (59940) must match the size of tensor b (1451808) at non-singleton dimension 2
I solve the problem,ComfyUI_smZNodes or ComfyUI-TiledDiffusion have some trouble with it.
@sekaiwlc07860 I ran into the same problem; how did you solve it? Thanks.
The samples already look great.. hope someone does an MXFP8 conversion at some point.
why do you want mxfp8 when int8 convrot is better quality and gives you the same speed?
@MrFlex If you have a 50 series card mxfp8 is faster than int8 convrot
@jervis314 i have a 50 series card and its not faster, and way worse quality, nvfp4 is faster but really bad quality
@MrFlex I have a 5090 and mxfp8 was definitely faster than int8 when I tested it with LTX and Krea 2. There was no difference in quality.
这里的int4到底可用吗?会不会质量稀碎?
Works on my 3060 12gb + 32gb ram using the native Load Diffusion Model node. ConvRot INT4 Layers mixed with INT8 Layers quantization to lower model size while maintaining quality over the speed benefit of INT4 as it results in poor quality.
I hope 4/8 step turbo lora realese soon
INT4BQ Balanced-Quality - what is it? Its faster or what? How about quality?
ConvRot INT4 Layers mixed with INT8 Layers quantization to lower model size while maintaining quality over the speed benefit of INT4 as it results in poor quality.
You will live to see man made horrors beyond your comprehension.
Lets hope so...
So I guess I need to install comfy now?
Yup
Has anyone else encountered the following error in ComfyUI?
# ComfyUI Error Report ## Error Details - Node ID: 128 - Node Type: CLIPLoader - Exception Type: AttributeError - Exception Message: AttributeError: 'NoneType' object has no attribute 'Params'
Appears when trying to load the text_encoder node and the main model node too.
I'm using the official workflow and the most recent version of ComfyUI.
LORAS for this model...IMMEDIATELY!
"light", UNCENSORED since base, easy use in confy, and i can make 768X768 (maybe higher)... with 8GB VRAM...nothing more to say except... i'm satisfied.
Toolkit already supports it: https://github.com/ostris/ai-toolkit/commit/8502a845b16418b1042dd7e3b2d819cc86326b1c
4060 8gb run ?
For which one of these 4 models are you referencing?
@BinaryBottleBake REF2VA INT8 pruned.
@Samoht yea
finally managed to run this model on my 3060 12gb and 16gb ram.. all you need is set pagefile.sys to at least 200gb+
generating 10 second takes 30 min at 0.4 megapixel. better than nothing lol
edit: use sage att + easy cache, now it's 60% faster
Strange, i use a 3060 12GB card and my gens were alot faster using the default workflow provided by comfy. I have since reworked it to use EasyCache and Sage and its faster, but 30 mins for a 10 second gen on a 3060 doesn't sound right.
I was generating at 0.7MP 16:9 Widescreen FirstFrame with default workflow at around 8-10 mins.
Are you using Int8/int4 or FP8?? An RTX 3060 12GB shouldn't take 30 minutes to generate a 0.4 MP image 10 seconds, int8.
Are you using atleast CUDA 13.0 (cu130) in your ComfyUI? Also, is the Comfy backend CUDA or Triton enabled , (disabled = false)? you can see Comfy backend info at the comfyui startup in your terminal.
If so, you should be getting the INT8/INT4 speedup. A 30 minute generation time on 12GB vram suggests that the CUDA or Triton backend isn't enabled (disabled = true), so the optimized kernels aren't being used.
Very impressive model! the prompt and camera adherence is remarkable. Only thing that's missing is some proper NSFW loras.
Indeed. I dont think i want to return to LTX now. I have so much fun playing with this one.
I'm gonna have to pull my rtx 6000pro from my llm box for this when porn lora are fleshed out
Waiting for mxfp8 safetensors version as I often find it fastest on my low RTX4060 8GB VRAM.
This is the model that we all have been waiting for. IMO, This + kREA 2 combo for I2V is going to be hard to top for foreseeable future.
The REF2VA has amazing image edit capabilities. just set frame lenght to 5 and select the best frame out of the results, works insanely fast even at large resolutions xD
MINIMAX H3 + KREA 2 MIGHT BE THE BEST AI MODEL YET TO GET EXACTLY WHAT U WANT
So good, MiniMax-H3 wipes the floor with LTX2.3
I have deleted all models of LTX2.3 from my pc
Do I need the text encoder and vae you put up to run the REF2VA-INT8 model?
We have one of the greatest achievement and progress in AI videography and the first thing people make are weird adult videos. I love it!!! :D
I removed all LTX 2.3 models and switched to MiniMax-H3
Wait for the lighten lora to decrease the step.
Now the step=20 need 4min to generate and the GPU temp is up to 77, and hot spot=92