MUST read务必读完
16g显存+32g内存可以跑1.0分辨率,56帧用时十分钟,没加vqa
使用千问VQA,1.0分辨率,6秒12帧,稳定960秒出图。
使用TE加速0.12,稳定620秒。
千问VQA,1.0分辨率,8秒12帧,TE加速0.12,17分钟。
我之前使用的是fp8版本,用int8用时是950秒
16GB VRAM + 32GB RAM can handle 1.0 resolution. 56 frames took 10 minutes without adding VQA.
Using QwenVQA (1.0 resolution, 6 seconds, 12 fps): Stable generation time of 960 seconds.
With TE acceleration at 0.12: Stable at 620 seconds.
QwenVQA (1.0 resolution, 8 seconds, 12 fps) + TE acceleration at 0.12: 17 minutes.
I was previously using the FP8 version, and using INT8 took 950 seconds.
minimax提示词不能写pussy,要写crotch,写pussy会画出来奇怪的东西
In Minimax prompts, don't use 'pussy'; use 'crotch' instead. Writing 'pussy' will make it generate weird things.
V 2.0
I have finally completed this workflow! By tweaking the prompts, I successfully managed to keep the model's breast size consistent. In the minimax视频生成改's node 205, I've prepared two sets of preset prompts: one for standard reference-based generation, and the other for anime-to-realistic conversion. However, you can still adjust specific details like skin texture to your own preferences.
As a bonus, I am also including my modified workflow for generating videos from the first and last frames. I added an acceleration LoRA node, but please note that this LoRA only works for 24 frames (so I don't actually use it myself), and it cannot be used simultaneously with TE acceleration. Furthermore, using an "Unload Model" node in the workflow will cause the acceleration LoRA to throw an error, which means this acceleration LoRA and Qwen VQA probably cannot be used together.
The acceleration LoRA is only compatible with the fl2va model. Here is the original link for the acceleration LoRA: https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
我终于彻底完成了这个工作流,通过修改提示词,成功实现了让模型保持胸部大小一致性,在minimax视频生成改的节点205,我准备了两组预设提示词,一组是用于正常参考生图的,另一组是用于二次元图片转真人的,但是具体的皮肤质感等细节,各位可以凭喜好再自行修改。
另外附赠我修改过的首尾帧生成视频工作流。我添加了加速Lora节点,但注意,这个lora只对24帧生效,所以我用不上,而且它和TE加速不能并用,并且在工作流中使用卸载模型节点会导致加速lora报错,这意味着加速Lora和千问VQA可能不能一起使用。加速lora只适用于fl2va模型,以下是加速lora原贴:https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
V1.8
新增TEspeed加速节点,下载地址:https://github.com/HELPMEEADICE/TE-Speed-MiniMaxH3-OSS
附件的model.py是TE加速需要的文件,替换至ComfyUI\comfy\ldm\minimax
但更重要的是,在gemini的帮助下我终于得到了VQA节点最完美的提示词,可以让VQA输出符合要求的多段式情节构思,限制字数和废话,并且有效避免复读bug,大幅提升稳定程度,缩短了用时。
Added the TEspeed acceleration node, download link: https://github.com/HELPMEEADICE/TE-Speed-MiniMaxH3-OSS
The attached model.py is required for TE acceleration. Please use it to replace the existing file in ComfyUI\comfy\ldm\minimax
But more importantly, with Gemini's help, I've finally nailed the absolute perfect prompt for the VQA node! It enables the VQA to output multi-paragraph, story-driven plot concepts that perfectly meet the requirements. It strictly limits the word count and cuts out the fluff, effectively avoids the repetition/infinite loop bug, massively boosts overall stability, and significantly reduces generation time.
v1.7
Added a frame extraction node, optimized the prompts and default settings, and included a bonus VQA testing workflow along with an image/video bidirectional conversion workflow.
新增加了帧提取节点,优化了提示词和默认设置,附赠VQA测试工作流和图像视频相互转换工作流
v1.5
Major Update V1.5: I really shouldn't have named the previous release V1.0—this is the true, complete, and official version.
Here are the key updates:
Heavily Optimized Default Prompts: Since I switched to the abliterated version of the Qwen3-VL-4B-FP8 model, the prompts can now be written much more directly and explicitly. I highly recommend everyone switch to the abliterated model as well. You can get it here: https://huggingface.co/huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated-FP8/tree/main
New Text Concatenation Logic: The workflow now utilizes two dual-segment text concatenation input boxes and introduces a new multi-segment text merge node.
Sequence-Aware Prompting: Because video generation heavily relies on frame sequence and timing, prompts are now required both before and after the VQA description.
Tweaked VQA Settings: I've fully optimized the VQA parameters. Unlike static image generation, video generation demands a significantly larger volume of text output from the vision model.
1.5重大更新,我之前不该命名为1.0的,现在才是正式完整版。
大幅优化了默认提示词设置,由于我更换了破甲版本的qwen3vl4bfp8模型,提示词可以写的更直接了,建议大家也使用破甲模型,地址:https://huggingface.co/huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated-FP8/tree/main
使用了两个两段文本拼接输入框,新加入了多段文本合并节点。视频涉及帧顺序,所以在VQA的描述前后都需要提示词。
优化了VQA的设置,它和生图不同,对文本量需求更大。
And then, heads up!
然后注意!
我花了一个晚上调整完了提示词,开始批量生成,然后睡了一觉,醒来发现一个8秒视频要用20分钟,qwenVQA就要用12分钟。现在它恢复正常了,但我不确定是什么原因导致的,以下是可能的原因:
1 我在昨天更换了破甲版本的qwen模型,
可能是更换模型后,minimax的页面没有刷新过(现在刷新过了),它在寻找旧的模型。这是最可能的原因,因为单独运行vqa时是正常的,更换模型前也是正常的,测试破甲模型时我关闭了VQA,直接输入文本来生图,所以没能发现跑1100秒的bug。
2 VQA温度0.8,输出上限200token,特定组合导致了bug。
3累计内存占用溢出,可能性比较小,因为它太突然了,而且minimax貌似并没有这种问题。保险起见我加入了清理节点。
接下来说说更新内容,1.0的时候,它无法再像简单提示词(jiggling breasts, exposing nipples)时那样出大特写。
我一开始尝试,让它像起初那样以特写开场,但是结果是如果让它第一个镜头出乳摇特写,然后第二镜头再正常出,它第一个镜头会默认是巨乳的。哦,天呐,巨乳就这么受欢迎吗。
所以我加了多段文本合并,通过前置VQA的反推,成功将特写镜头改到了最后。
最终结果,我默认设置是6-8秒的视频。
另外minimax的阴毛效果不好,尤其是对阴毛的位置理解,只有正面视角,不会露出屁股的时候,才适合添加阴毛。
另外,感觉minimax真人效果比动漫好很多。
I spent the whole night tweaking the prompts and started a batch generation. Then I took a nap, only to wake up and find that a single 8-second video was taking 20 minutes, with the Qwen VQA alone taking 12 minutes! It's back to normal now, but I'm not entirely sure what caused it. Here are the possible reasons:
I switched to the abliterated version of the Qwen model yesterday. It's possible that after switching models, the Minimax page hadn't been refreshed (it is refreshed now), so it was still searching for the old model. This is the most likely reason because running the VQA in isolation works fine, and everything was normal before the model swap. When I tested the abliterated model, I actually disabled the VQA and inputted text directly to generate images, which is why I didn't catch this 1100-second bug earlier.
The VQA temperature was set to 0.8 with a max output limit of 200 tokens. This specific combination might have triggered a bug.
Cumulative memory overflow. This is less likely because it happened too suddenly, and Minimax doesn't seem to suffer from this issue. Just to be safe, though, I added a cleanup node.
Now, let's talk about the updates. In version 1.0, it could no longer pull off the extreme close-ups like it used to with simple prompts (like "jiggling breasts, exposing nipples").
At first, I tried to have it open with a close-up like I did originally. But the result was that if I made the first shot a breast-jiggling close-up and the second shot a normal angle, it would default to giving her massive breasts in the first shot. Oh my god, are huge breasts really that popular?
So, I added a multi-segment text merge. By utilizing the front-end VQA's reverse deduction, I successfully moved the extreme close-up shot to the very end.
The final result is that my default setting is now a 6-8 second video.
Additionally, Minimax doesn't handle pubic hair very well, especially when it comes to understanding its placement. It's really only suitable to add pubic hair when using a frontal perspective where the buttocks aren't exposed.
Also, I feel like Minimax's realistic/live-action generation is significantly better than its anime generation.
V1.0
As always, we kick things off with the Batch Image node, followed by JoyTag and QwenVQA, keeping everything fully automated.
I didn't originally want to make it this complicated, but the workflows for the first and last frames both rely on RunningHub nodes.
The Qwen3-VL-4B fp8 model I'm using is quite stubborn and isn't uncensored (jailbroken). The standard uncensored Qwen3-VL-4B combined with 8-bit quantization is just way too slow. As a result, the prompt combinations might look a bit weird.
The default settings are a 2:3 aspect ratio, 720p resolution, and 12 fps. Important note: The frame count must be modified in Node 131, NOT Node 130.
Note: The SDPA acceleration used by the VQA node will freeze/crash under xformers.
JoyTag probably doesn't need to be enabled—you can decide whether to turn it on based on your specific needs.
依旧批量图片节点起手,依旧joytag起手,依旧qwenvqa起手,依旧自动化。
我本来不想搞这么复杂的,但是首尾帧的工作流都是用的runinghub的节点。
我用的QWEN3vl4b的fp8模型比较顽固,没有破甲。破了甲的普通QWEN3vl4b配合8bit压缩太慢了。所以提示词的组合比较怪异。
默认设置是2:3,720p,每秒12帧,注意注意,帧数在节点131修改而不是130。
注意:VQA节点用的sdpa加速在xformers下会卡死。
joytagy应该不需要打开,可根据需要决定。








