CivArchive
    [MiniMax H3] NSFW I2VA / T2VA / FL2VA / R2VA Workflows ๐Ÿ‘พ Qwen3VL Auto Prompt | Native Audio | TensorRT Upscale | RIFE Interpolation - ๐ŸŽฅ OneClick-T2VA
    NSFW

    ComfyUI-QwenVL-Mod

    Enhanced Vision-Language with MiniMax H3 Version 2.5 (2026/08/05)

    ๐ŸŽฌ MiniMax H3 Native Video+Audio + Qwen3-VL Auto-Prompting

    ๐ŸŸฃ Deploy on Runpod

    ๐ŸŸก Deploy on Vast.ai


    โš ๏ธ Requirements โ€” Read First!

    GPU & VRAM

    • ๐ŸŸข Recommended โ€” RTX 5090 / 4090 (24 GB+) โ†’ INT8 pruned โ†’ Fast, best quality

    • ๐ŸŸก Enthusiast โ€” RTX 3090 / 4080 (16-24 GB) โ†’ INT8 pruned + offload โ†’ Slower, good quality

    • ๐ŸŸ  Budget โ€” RTX 3060 / 4060 Ti (12-16 GB) โ†’ INT4 pruned + offload โ†’ Slow, usable quality

    • ๐Ÿ”ด Experimental โ€” Blackwell GPUs (12 GB+) โ†’ NVFP4 โ†’ Requires Blackwell Tensor Cores

    12 GB GPUs (e.g. RTX 3060 12GB): Technically possible with INT4 models + aggressive offloading, but very slow. You need 32 GB+ system RAM and a fast NVMe SSD. Not recommended for production use.

    Model Quantization Options

    Software

    • ComfyUI: v0.30.0+ (required for MiniMax H3 native support)

    • Python: 3.10+

    • CUDA: 12.8+ (13.0 recommended)

    • Storage: 30-110 GB SSD depending on quantization

    Qwen3-VL Prompt Enhancer

    • GGUF: Q4_K_S (~4.8 GB) or Q5_K_S (~5.5 GB) for 8B model

    • HF: Qwen3-VL-8B-Heretic-Stable (~16 GB) or Qwen3-VL-4B (~8 GB)


    ๐ŸŒŸ What is ComfyUI-QwenVL-Mod?

    A powerful enhanced vision-language node for ComfyUI that combines Qwen3-VL models with MiniMax H3 video generation workflows. Features multilingual support, visual style detection, native stereo audio, and NSFW capabilities for professional AI content creation.

    Think: "Your all-in-one solution for intelligent prompt enhancement and video+audio generation with MiniMax H3!"


    ๐ŸŽฌ Key Features

    ๐Ÿš€ MiniMax H3 Video+Audio Generation

    • T2VA (Text-to-Video+Audio): Generate video with native stereo audio from text

    • I2VA (Image-to-Video+Audio): Animate a first-frame image with audio

    • FL2VA (First-Last-Frame): Generate the transition between two keyframes โ€” Qwen3-VL sees both frames

    • R2VA (Reference-to-Video): Lock character identity, style, motion, or voice using reference images

    ๐Ÿง  Qwen3-VL Auto-Prompting

    • Multilingual: Write your prompt in any language โ€” Qwen3-VL translates and converts it

    • Auto-format: Generates the official MiniMax H3 prompt format (3-field for base, 6-field for R2VA)

    • Multi-reference: Qwen3-VL sees all connected images via image + image2 inputs

    • Visual style detection: 12+ artistic styles (photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy, etc.)

    • Smart caching: Performance optimization with Fixed Seed Mode

    • GGUF backend: Efficient local model inference with quantization support

    • Qwen3.5 support: Thinking mode disabled via /no_think for fast prompt generation

    ๐Ÿ”Š Native Stereo Audio

    • No separate audio node needed โ€” MiniMax H3 generates video and audio jointly in a single forward pass

    • Voice, sound effects, and music modeled together, not layered on afterward

    • Describe sounds in your prompt and the model generates them natively

    ๐ŸŽจ NSFW Support

    • Comprehensive content generation without restrictions

    • 9 dedicated NSFW presets (3 base ๐ŸŽฌ + 3 R2VA ๐ŸŽž๏ธ + 3 FL2VA ๐Ÿ”„) with explicit diegetic soundscape

    • Natural progression, style adaptation, consistent characters


    ๐Ÿ“ฆ What's Included โ€” 4 Workflows

    1. ๐Ÿ“ T2VA โ€” MiniMaxH3-T2VA-Qwen3VL.json โ€” text only โ€” Text-to-video+audio. Simplest workflow. Uses PromptEnhancer (text-only).

    2. ๐Ÿ–ผ๏ธ I2VA โ€” MiniMaxH3-I2VA-Qwen3VL.json โ€” text + first-frame image (image) โ€” Image-to-video. First-frame animation with audio.

    3. ๐Ÿ”„ FL2VA โ€” MiniMaxH3-FL2VA-Qwen3VL.json โ€” text + first-frame (image) + last-frame (image2) โ€” First-Last-Frame to video. Qwen3-VL sees both frames and describes the transition. Includes TensorRT upscale + RIFE frame interpolation for 48 fps output.

    4. ๐ŸŽž๏ธ R2VA โ€” MiniMaxH3-R2VA-Qwen3VL.json โ€” text + reference images (image + image2) โ€” Reference-to-video. Qwen3-VL sees all references. Lock identity, style, motion, camera, or voice using up to 9 ref images.

    Workflows 3 and 4 include TensorRT upscaling (RealESRGAN x4) and RIFE frame interpolation (rife49) for 48 fps high-resolution output.


    ๐Ÿ–ผ๏ธ Multi-Reference Input (image2)

    The QwenVL-Mod node has two image inputs:

    • T2VA: no images needed

    • I2VA: image = first frame

    • FL2VA: image = first frame, image2 = last frame, frame_count = 1

    • R2VA: image = primary reference, image2 = additional references (batch, up to 9), frame_count = 1โ€“9

    Qwen3-VL sees all connected images as individual images (not as a video sequence), enabling proper multi-reference analysis for FL2VA and R2VA.


    ๐ŸŽฏ QwenVL-Mod NSFW Presets (9 total)

    The workflows include built-in NSFW presets for the Qwen3-VL prompt enhancer:

    ๐ŸŽฌ Base Presets (T2VA / I2VA)

    • ๐ŸŽฌ MiniMax H3 NSFW (5s) โ€” 5 seconds โ€” 3 fields: integrated_multimodal_description + overall_soundscape + non_diegetic_music

    • ๐ŸŽฌ MiniMax H3 NSFW (10s) โ€” 10 seconds โ€” Same format

    • ๐ŸŽฌ MiniMax H3 NSFW (15s) โ€” 15 seconds โ€” Same format

    ๐Ÿ”„ FL2VA Presets (First-Last-Frame)

    • ๐Ÿ”„ MiniMax H3 NSFW FL2VA (5s) โ€” 5 seconds โ€” 3 fields, transition-focused (describes the path between frames)

    • ๐Ÿ”„ MiniMax H3 NSFW FL2VA (10s) โ€” 10 seconds โ€” Same format

    • ๐Ÿ”„ MiniMax H3 NSFW FL2VA (15s) โ€” 15 seconds โ€” Same format

    ๐ŸŽž๏ธ R2VA Presets (Reference)

    • ๐ŸŽž๏ธ MiniMax H3 NSFW R2VA (5s) โ€” 5 seconds โ€” 6 fields: subject_definitions + summary + retention_analysis + detailed_description + overall_soundscape + non_diegetic_music

    • ๐ŸŽž๏ธ MiniMax H3 NSFW R2VA (10s) โ€” 10 seconds โ€” Same format

    • ๐ŸŽž๏ธ MiniMax H3 NSFW R2VA (15s) โ€” 15 seconds โ€” Same format

    What the presets produce

    • ๐ŸŽฌ Base: [Shot 1] with style + initial composition, camera vocabulary, speaker IDs, diegetic soundscape

    • ๐Ÿ”„ FL2VA: Describes the transition path between first and last frames (not the scene โ€” images fix the scene). Favors single continuous shot.

    • ๐ŸŽž๏ธ R2VA: 6-section format with <Subject N>, <Picture N>, <Video N>, <Audio N> labels, retention markers (fully_preserved, partially_preserved, etc.), task-type summary

    • All presets: smooth, continuous camera motion (no abrupt or stepped changes), explicit diegetic soundscape, optional non-diegetic music (defaults to N/A)

    SFW presets are also available. Edit the preset dropdown in the QwenVL node to switch.


    ๐ŸŽฎ Usage Examples

    Basic Text-to-Video (T2VA)

    1. Load MiniMaxH3-T2VA-Qwen3VL.json

    2. Write your prompt in any language

    3. Select preset ๐ŸŽฌ MiniMax H3 NSFW (5s/10s/15s)

    4. Generate video with native audio

    Image-to-Video (I2VA)

    1. Load MiniMaxH3-I2VA-Qwen3VL.json

    2. Upload your first-frame image to image

    3. Select preset ๐ŸŽฌ MiniMax H3 NSFW (5s/10s/15s)

    4. Write what happens next (in any language)

    5. Generate animated video with audio

    First-Last-Frame (FL2VA)

    1. Load MiniMaxH3-FL2VA-Qwen3VL.json

    2. Upload first-frame to image, last-frame to image2, set frame_count=1

    3. Select preset ๐Ÿ”„ MiniMax H3 NSFW FL2VA (5s/10s/15s)

    4. Describe the transition between the two frames

    5. Generate the interpolated video at 48 fps with TensorRT upscale + RIFE

    Reference-to-Video (R2VA)

    1. Load MiniMaxH3-R2VA-Qwen3VL.json

    2. Upload primary reference to image, additional references to image2 (batch), set frame_count to match

    3. Select preset ๐ŸŽž๏ธ MiniMax H3 NSFW R2VA (5s/10s/15s)

    4. Reference them by tag in your prompt: <Picture 1>, <Picture 2>, etc.

    5. Generate video with locked identity/style


    ๐Ÿ”ง Technical Specifications

    โšก Performance

    • Output: 768p, 24 fps (native), up to ~15 seconds

    • Audio: Native stereo, generated jointly with video

    • Upscale: TensorRT RealESRGAN x4 (FL2VA + R2VA workflows)

    • Frame interpolation: RIFE rife49 โ†’ 48 fps (FL2VA + R2VA workflows)

    • Sage Attention: FP16 accumulation, async offload

    • Smart caching: Reuse prompts with same inputs, Fixed Seed Mode for text-only caching

    ๐ŸŽจ Model Support

    • Qwen3-VL 4B: 7 GGUF variants (2.38 GB โ€“ 4.28 GB)

    • Qwen3-VL 8B: 7 GGUF variants (4.8 GB โ€“ 8.71 GB)

    • Qwen3.5: 4B / 9B / 27B (uncensored, heretic, unsloth) โ€” thinking mode disabled

    • HF Models: Josiefed, official, Heretic-Stable variants

    • Quantization: Q4_K_S, Q5_K_S, FP16, INT8

    ๐ŸŒ Multilingual Capabilities

    • Input languages: Any language supported

    • Auto-translation: Automatic translation to optimized English

    • Style detection: Works with multilingual prompts

    • Cultural adaptation: Context-aware prompt enhancement


    ๐Ÿ“ฆ Installation

    Quick Install

    1. Download: ComfyUI-QwenVL-Mod (latest version)

    2. Extract to ComfyUI/custom_nodes/ComfyUI-QwenVL-Mod

    3. Install requirements: pip install -r requirements.txt

    4. Restart ComfyUI

    5. Load included workflows from minimax/ folder

    Custom Nodes Required

    Models Required

    All MiniMax H3 models from Comfy-Org/MiniMax-H3 on Hugging Face.

    T2VA / I2VA / FL2VA (fl2va)

    • models/vae/ โ†’ minimax_h3_video_vae_fp16.safetensors (~5 GB)

    • models/vae/ โ†’ minimax_h3_audio_vae_fp32.safetensors (~0.6 GB)

    • models/diffusion_models/ โ†’ minimax_h3_fl2va_pruned_int8_convrot.safetensors (~21 GB)

    • models/text_encoders/ โ†’ qwen3vl_32b_minimax_h3_int8_convrot.safetensors (~27 GB)

    R2VA (ref2va) โ€” same as above, except:

    • models/diffusion_models/ โ†’ minimax_h3_ref2va_pruned_int8_convrot.safetensors (~21 GB)

    INT4 alternative (for 12-16 GB GPUs): Merserk/MiniMax-H3-INT4-ConvRot

    Qwen3-VL Prompt Enhancer

    • models/LLM/ โ†’ Qwen3-VL-8B-Heretic-Stable (GGUF or HF)

    TensorRT Engines (FL2VA + R2VA only)

    • models/upscale_models/ โ†’ RealESRGAN_x4 (TensorRT engine)

    • models/rife/ โ†’ rife49_ensemble_True_scale_1_sim (TensorRT engine)

    TensorRT engines must be built for your specific GPU. See ComfyUI-RIFE-TensorRT-Auto and ComfyUI-Upscaler-TensorRT-Auto for build instructions.


    ๐ŸŽฌ MiniMax H3 Prompting Notes

    How to Write Your Prompt

    Describe the scene naturally. Be clear about the concepts below โ€” Qwen3-VL handles the rest:

    • ๐ŸŽจ Visual style (put it first): photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy

    • ๐Ÿ‘ฅ Subjects: number, gender, appearance, clothing, position, expression

    • ๐Ÿƒ Action / motion: what happens, speed, interaction

    • ๐ŸŽฅ Camera: dolly, pan, zoom, static, handheld, crane, orbit โ€” smooth and continuous (no abrupt changes)

    • ๐ŸŒ Environment: setting, lighting, atmosphere, time of day

    • ๐Ÿ”Š Audio (important!): dialogue, breaths, moans, skin contact, ambient sounds, music

    ๐Ÿ”„ FL2VA: Describe the transition between frames, not the scene (images fix the scene) ๐ŸŽž๏ธ R2VA: Reference inputs by tag: <Picture 1>, <Picture 2>, <Video 1>, <Audio 1>

    Resolution Guidance

    MiniMax H3 native canvas: 768 px short edge, long edge capped at 1344 px, multiples of 32.

    • ๐Ÿ“ฑ Portrait: 768ร—1344 ยท 896ร—1152 ยท 960ร—1280

    • โฌ› Square: 1024ร—1024

    • ๐Ÿ–ฅ๏ธ Landscape: 1344ร—768 ยท 1152ร—896 ยท 1280ร—960

    โš ๏ธ Match the aspect ratio to your input image! Forcing 16:9 on a portrait image will squash it.

    โš ๏ธ Avoid direct 1080p. Generate at native resolution, then upscale with TensorRT nodes (FL2VA + R2VA workflows).

    Duration

    Choose a preset: 5s / 10s / 15s. The Math Expression node snaps the frame count to the model's 17-frame-per-block grid (17k+5 at 24 fps).


    ๐Ÿณ Docker / Cloud Ready

    OneClick RunPod Template

    Prefer a ready-to-go environment? Use the OneClick ComfyUI MiniMax H3 Qwen3VL RunPod template:

    • Docker image: huchukato/comfyui-qwenvl-runpod:cu13-minimax

    • Base: runpod/comfyui:cuda13.0

    • All custom nodes pre-installed

    • All 4 workflows auto-downloaded at boot

    • Models auto-downloaded at first boot (~50 GB, persistent)

    • ComfyUI v0.30.0+ forced at boot

    • Sage Attention, FP16 accumulation, async offload

    • TensorRT upscaling + RIFE interpolation

    ๐Ÿ“– README & instructions

    ComfyUI Args (pre-configured)

    --disable-auto-launch
    --fast fp16_accumulation
    --use-sage-attention
    --reserve-vram 2
    --cuda-malloc
    --async-offload
    

    ๐Ÿš€ Why Choose ComfyUI-QwenVL-Mod + MiniMax H3?

    ๐ŸŽฌ For Content Creators

    • Native audio: Video and audio in one pass โ€” no separate MMAudio needed

    • Multilingual: Write in any language, Qwen3-VL handles translation

    • Professional: Official MiniMax H3 prompt format with camera vocabulary and speaker tags

    • Quality: 768p native, TensorRT upscale to higher resolution

    ๐Ÿ”ฅ For NSFW Content

    • Explicit: Uncensored generation with dedicated NSFW presets

    • 9 presets: 3 base ๐ŸŽฌ + 3 FL2VA ๐Ÿ”„ + 3 R2VA ๐ŸŽž๏ธ โ€” each tuned for its mode

    • Detailed: Rich scene descriptions with explicit diegetic soundscape

    • Natural: Realistic progression, consistent characters

    • Audio: Native moans, breaths, skin contact, ambient sounds

    โšก For Power Users

    • Customizable: Easy to modify presets and system prompts

    • Extendable: Add your own Qwen3-VL models (GGUF or HF)

    • Integrable: Works with existing ComfyUI setups

    • Optimized: Sage Attention, FP16, async offload, smart caching

    • Multi-reference: image2 input for FL2VA and R2VA workflows


    ๐ŸŒŸ What Makes This Special?

    • First: Complete MiniMax H3 workflow pack with Qwen3-VL auto-prompting

    • Native audio: No separate audio node โ€” MiniMax H3 does it all

    • 4 workflows: T2VA, I2VA, FL2VA, R2VA โ€” covers all MiniMax H3 modes

    • Multi-reference: Qwen3-VL sees all connected images (not just the first)

    • TensorRT: Built-in upscaling and frame interpolation

    • 9 NSFW presets: Dedicated presets for each mode with correct prompt structure

    • Multilingual: Any input language, auto-translated and formatted

    • Ready: Works out-of-the-box with included workflows


    ๐ŸŽฏ What's New in v2.5

    ๐Ÿš€ MiniMax H3 Full Support

    • โœ… 4 workflows: T2VA, I2VA, FL2VA, R2VA โ€” all modes covered

    • โœ… Multi-reference input: image2 input โ€” Qwen3-VL sees all images as individual images

    • โœ… 9 NSFW presets: 3 base ๐ŸŽฌ + 3 R2VA ๐ŸŽž๏ธ + 3 FL2VA ๐Ÿ”„ with correct prompt structure

    • โœ… Smooth camera: All presets enforce smooth, continuous camera motion

    • โœ… Native audio: Video + stereo audio in one pass

    • โœ… Official format: 3-field (base) and 6-field (R2VA) prompt formats

    ๏ฟฝ Qwen3.5 Thinking Fix

    • โœ… /no_think prefix for Qwen3.5 models (enable_thinking deprecated in recent llama.cpp)

    • โœ… Broadened architecture detection (qwen35, qwen35moe, qwen35_vl)

    • โœ… Works across both HF and GGUF nodes

    ๐Ÿ“ฆ Workflow Organization

    • โœ… Moved workflows to minimax/ folder

    • โœ… Renamed FLF to FL2VA (clearer naming)

    • โœ… Added Civitai documentation


    ๐Ÿ“‹ Credits


    ๐Ÿ“„ License

    Workflows are released under the same license as the underlying models and custom nodes. See each repository for details.

    MiniMax H3 model weights: Comfy-Org/MiniMax-H3 โ€” MiniMax H3 Community License.


    Built with โค๏ธ for the ComfyUI community

    Description

    FAQ

    Workflows
    MiniMax H3

    Details

    Downloads
    84
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/5/2026
    Updated
    8/6/2026
    Deleted
    -

    Files

    MinimaxH3NSFWI2VAT2VAFL2VAR2VA_OneclickT2VA.zip

    MinimaxH3NSFWI2VAT2VAFL2VAR2VA_OneclickT2VA.zip