What this is?
My H3 workflow, designed to have quick generations on my relatively low VRAM (12gb 3060).
On my setup, this workflow has the following generation times (compared to a basic 4 step workflow), rounded to the nearest 30 seconds:
5 seconds, I2V
| Megapixels: | Basic 4 Step | V1.0 |
---------------------------------------
| 0.5 | 3:30 | 2:30 |
| 1.0 | 9:30 | 5:30 |
| 2.0 | 30:00 | 13:30 |
10 seconds, I2V
| Megapixels: | Basic 4 Step | V1.0 |
---------------------------------------
| 0.5 | 9:00 | 4:30 |
| 1.0 | 27:30 | 10:00 |
| 2.0 | OOM, DNF | 20:30 |Features
Supports both ref2va and fl2va
Select a fl2va model and toggle the 'use fl2va?' option to run in fl2va mode, leave it off and load a ref2va model for ref2va mode.
Automatically constructs prompts
Each section of a H3 prompt (subject_descriptions, retention_analysis, etc.) has it's own input, all of which are combined to construct the prompt, making it easier to keep track of everything.
Dynamic prompts
Uses the dynamic prompts extension to do some pretty neat things, the most useful in this case is using variables in your prompt.
LoRA Support
Supports any LoRAs you want. Note: Adding LoRAs will increase generation time.
Easy references
Supports the maximum amount of references ref2va uses (9 image, 3 video and 3 audio), and will automatically disable image references 9, 8 and 7 for each reference over the maximum of 12.
Resize references
Allows the resizing of image and video references to match the selected generation size, useful for fl2va and keyframes.
Required node packs
Uses the following node packs:
🟢 Will be auto-detected by comfyui-manager
🟡 Can be installed using comfyui-manager, but is not auto-detected (you will have to search for it). Manual installation may be easier to avoid getting the wrong pack.
🔴 Requires manual installation🟢 KJ Nodes
🔴 MiniMax H3 Audio T8 Note: This one can be installed using comfyui-manager (it is not auto-detected), but reports it's compatibility requirements incorrectly so it cannot easily be enabled. Manual installation bypasses that.
Required models
Diffusion model
Any int8-convrot model, I use these hybrid models for ref2va, and the basic fl2va one. These go in models/diffusion_models
Text encoder
This one. Goes in models/text_encoders
Vae
Video, audio. Goes in models/vae
Latent upscaler
This one. Goes in models/latent_upscale_models
4 Step LoRA
Any pruned LoRA model (must be pruned, will crash otherwise). I use this one, this one is also pretty good. Goes in models/loras
Full model paths
📂 ComfyUI
├── 📂 models
| ├── 📂 diffusion_models
| | ├── DIFFUSION MODELS GO HERE
| |
| ├── 📂 latent_upscale_models
| | ├── LATENT UPSCALER GOES HERE
| |
| ├── 📂 loras
| | ├── 4 STEP LORA GOES HERE
| |
| ├── 📂 text_encoders
| | ├── TEXT ENCODER GOES HERE
| |
| ├── 📂 vae
| | ├── VAE MODELS GO HEREDescription
First version I'm happy with. I have a few more things I'd like to add further down the track, but this is a pretty good starting point.
Features
Supports both ref2va and fl2va
Select a fl2va model and toggle the 'use fl2va?' option to run in fl2va mode, leave it off and load a ref2va model for ref2va mode.
Automatically constructs prompts
Each section of a H3 prompt (subject_descriptions, retention_analysis, etc.) has it's own input, all of which are combined to construct the prompt, making it easier to keep track of everything.
Dynamic prompts
Uses the dynamic prompts extension to do some pretty neat things, the most useful in this case is using variables in your prompt.
LoRA Support
Supports any LoRAs you want. Note: Adding LoRAs will increase generation time.
Easy references
Supports the maximum amount of references ref2va uses (9 image, 3 video and 3 audio), and will automatically disable image references 9, 8 and 7 for each reference over the maximum of 12.
Resize references
Allows the resizing of image and video references to match the selected generation size, useful for fl2va and keyframes.
Required node packs
Uses the following node packs:
🟢 Will be auto-detected by comfyui-manager
🟡 Can be installed using comfyui-manager, but is not auto-detected (you will have to search for it). Manual installation may be easier to avoid getting the wrong pack.
🔴 Requires manual installation🟢 KJ Nodes
🔴 MiniMax H3 Audio T8 Note: This one can be installed using comfyui-manager (it is not auto-detected), but reports it's compatibility requirements incorrectly so it cannot easily be enabled. Manual installation bypasses that.
Required models
Diffusion model
Any int8-convrot model, I use these hybrid models for ref2va, and the basic fl2va one. These go in models/diffusion_models
Text encoder
This one. Goes in models/text_encoders
Vae
Video, audio. Goes in models/vae
Latent upscaler
This one. Goes in models/latent_upscale_models
4 Step LoRA
Any pruned LoRA model (must be pruned, will crash otherwise). I use this one, this one is also pretty good. Goes in models/loras
Full model paths
📂 ComfyUI
├── 📂 models
| ├── 📂 diffusion_models
| | ├── DIFFUSION MODELS GO HERE
| |
| ├── 📂 latent_upscale_models
| | ├── LATENT UPSCALER GOES HERE
| |
| ├── 📂 loras
| | ├── 4 STEP LORA GOES HERE
| |
| ├── 📂 text_encoders
| | ├── TEXT ENCODER GOES HERE
| |
| ├── 📂 vae
| | ├── VAE MODELS GO HEREFAQ
Comments (2)
This WF works great. It runs quickly and has nice output. Thanks.
Thanks! I'm glad you like it.
