H3 turns references into hard-hitting action videos with native audio.
Who it's for: creators who want this pipeline in ComfyUI without assembling nodes from scratch. Not for: one-click results with zero tuning - you still choose inputs, prompts, and settings.
Open preloaded workflow on RunComfy
Open preloaded workflow on RunComfy (browser)
Why RunComfy first
- Fewer missing-node surprises - run the graph in a managed environment before you mirror it locally.
- Quick GPU tryout - useful if your local VRAM or install time is the bottleneck.
- Matches the published JSON - the zip follows the same runnable workflow you can open on RunComfy.
When downloading for local ComfyUI makes sense - you want full control over models on disk, batch scripting, or offline runs.
How to use (local ComfyUI)
1. Load inputs (images/video/audio) in the marked loader nodes.
2. Set prompts, resolution, and seeds; start with a short test run.
3. Export from the Save / Write nodes shown in the graph.
Expectations - First run may pull large weights; cloud runs may require a free RunComfy account.
Overview
Turn two character images and a scene reference into action video. You get native audio and identity-guided characters. Structure prompts to control choreography. Two-stage sampling refines motion. Build sword, boxing, or sci-fi combat clips. Review fast hand and weapon contact for artifacts.
Important nodes:
Key nodes in Comfyui MiniMax H3 Action Scenes workflow
MiniMaxH3ReferenceToVideo (#56)
Builds a synchronized audio‑video latent from your structured prompt and reference images using the MiniMax H3 text encoder and VAEs. Adjust the prompt, width, height, and length to control content, framing, and duration. For best results, keep the subject definitions and scene constraints unambiguous. Reference: Comfy‑Org MiniMax‑H3.
MinimaxH3LatentUpscaler3D (#124)
Temporally consistent 3D latent upscaler that enlarges the video latent between passes before refinement. Increase mode.scale modestly to gain detail without destabilizing motion. Useful when you like the pass‑1 timing but want sharper edges and textures. Reference: Comfyui Minimax H3 Latent Upscaler.
MiniMaxH3SigmaShift (#297)
Shifts the audio and video sigma schedules used by MiniMax H3 so motion, lip cues, and foley stay aligned. Tuning shift_video and shift_audio can reduce timing drift in dialogue or impacts. Leave small offsets unless you notice desync at cut‑in or cut‑out. Reference: MiniMax‑H3 collection.
VHS_VideoCombine (#264)
Combines decoded frames and audio into the final MP4 with your chosen frame rate and quality settings. Enable metadata saving to retain provenance and workflow details in exports. Use trim_to_audio when your refined video extends beyond the preserved audio. Reference: ComfyUI‑VideoHelperSuite.
PathchSageAttentionKJ (#199)
Attention optimization that can improve throughput and stability for high‑detail scenes and larger resolutions. Keep it enabled for longer takes or complex environments. If you encounter performance regressions on smaller clips, try its automatic mode. Reference: ComfyUI‑KJNodes.
Notes
MiniMax H3 Action Scenes | Reference Video + Native Audio - see RunComfy page for the latest node requirements.
Description
Comments (2)
It probably wouldn't hurt if you guys at least mentioned the system specs you used to test the workflows. In my case... 16 GB GPU, 32 GB system RAM, and I got a "not enough GPU memory" error.
And i chanded your 32 Gb main model with mine 20 Gb model, which works on most of my workflows.
