Base MinimaxH3 (Turbo Merge)
This model is unchanged/untrained from the base model listed. The model was quantized from the source FP32/BF16 model. BOTH THE FL2VA and REF2VA models are 8 step models.
Requirements:
Update CUDA to 13.4.2
Update Pytorch to 13.2
Update at minimum comfy-aimdo, comfy-kitchen
Basic workflow shows how to use both first frame text guided and first frame last frame with comfy kitchen backend node to speed up generation by 60% (This prevents pyattention fallback which is slow)
pytorch version: 2.13.0+cu132
xformers version: 0.0.35
Using xformers attention
ComfyUI version: 0.35.0
comfy-aimdo version: 0.5.5
comfy-kitchen version: 0.2.34
Consider --disable-dynamic-vram if you are having OOM issues after a few generation or crash when trying to use comfy kitchen vs pyattention
Note: INT Convrot Group size is 64 vs 256 on the comfy supplied models
NF4 version of the reference to video/audio model works but the quality is very low
Description
FAQ
Comments (8)
Maybe you should precise that Torchaudio is unavailable for cu132. You need to look for the info, "normal" install will just result to a fail next time you use your ComfyUI. Here if anyone update : https://huggingface.co/ussoewwin/torchaudio-built-on-cu132/tree/main
you're welcome :)
I have a working version of both torchao torchao-0.18.0 and torch audio torchaudio-2.11.0 - I have no issues with
TorchAudio 2.11 works with torch 2.11 and with every future torch release (2.12, 2.13, etc.). It is the latest release, and it is the only one you need: there is nothing to upgrade in TorchAudio when you upgrade torch, and installing TorchAudio does not pin torch to a specific version.
Nothing against compiling with the latest version, or posting them, I have done it myself with bitsandbytes but I don't think it is mandatory
@Felldude i think your misunderstand my point but that's ok. i know by saying "Update PyTorch to 13.2" people are going to have this issue. but hey, i was just saying ! :P
What's the purpose of merging an accelerator lora with a standard model? Why not just use a lora loader for the accelerator? That way different accelerator loras can be swapped in and the accelerator can be disabled on demand.
I'm haven't been able to think of any reason to do the merge other than to shrink the total size because the accelerator is displacing some of the base model's content. If size shrinkage is the goal I would think quantization would be the better route.
The lora was projected without quantization, then quantized, rather then the lora being applied to the upcast blocks as they are loaded/offloaded - It is also not the same turbo, REF is 4 step turbo I did not like that quality
这个模型和官方发布的区别在哪?谢谢