๐ฃ Deploy on Runpod

๐ก Deploy on Vast.ai

ComfyUI-QwenVL-Mod โ Enhanced Vision-Language with LTX 2.3 Version 2.3.1 (2026/08/25) โ ๐ฌ LTX 2.3 Native Video+Audio + 10Eros Uncensored + Qwen3.5 Auto-Prompting
โ ๏ธ Requirements โ Read First!
GPU & VRAM
๐ข Recommended โ RTX 5090 / 4090 (24 GB+) โ FP8 text encoder + checkpoint FP8 โ Fast, best quality
๐ก Enthusiast โ RTX 3090 / 4080 (16-24 GB) โ offload text encoder / VAE โ Slower, good quality
๐ Budget โ RTX 4060 Ti 16 GB โ aggressive offload + lower resolution โ Experimental, slow
12 GB GPUs: Not recommended for LTX 2.3 with the full 10Eros setup. The text encoder alone is ~12.8 GB FP8.
Software
ComfyUI: v0.32.0+ (required for LTX 2.3 native audio and Sage Attention flags)
Python: 3.10+
CUDA: 13.0 recommended
Storage: 100 GB+ SSD (models ~40 GB + outputs)
Qwen3-VL Prompt Enhancer
GGUF: Q4_K_S (~4.8 GB) or Q5_K_S (~5.5 GB) for 8B model
HF: Qwen3-VL-8B-Heretic-Stable (~16 GB) or Qwen3-VL-4B (~8 GB)
๐ What is ComfyUI-QwenVL-Mod?
A powerful enhanced vision-language node for ComfyUI that combines Qwen3-VL models with LTX 2.3 video generation workflows. Features multilingual support, visual style detection, native synchronized audio, and NSFW capabilities for professional AI content creation.
Think: *"Your all-in-one solution for intelligent prompt enhancement and uncensored video+audio generation with LTX 2.3!"
๐ฌ Key Features
๐ LTX 2.3 Native Video+Audio Generation
I2VA (Image-to-Video+Audio): Animate a first-frame image with synchronized audio
FL2VA (First-Last-Frame): Generate the transition between two keyframes with audio โ Qwen3-VL sees both frames
LTX 2.3 produces video and audio in the same forward pass: no separate MMAudio or audio-VAE stitching is required.
๐ง Qwen3.5 Auto-Prompting
Multilingual: Write your prompt in any language โ Qwen3-VL translates and converts it
Auto-format: Generates optimized LTX 2.3 prompt structure
Multi-reference: Qwen3-VL sees all connected images via
image+image2inputsVisual style detection: 12+ artistic styles (photorealistic, cinematic, anime, 3D CG, claymation, vintage film, watercolor, fantasy, etc.)
Smart caching: Performance optimization with Fixed Seed Mode
GGUF backend: Efficient local model inference with quantization support
Qwen3.5 support: Thinking mode disabled via
/no_thinkfor fast prompt generation
๐ Native Synchronized Audio
No separate audio node needed โ LTX 2.3 generates video and audio jointly
Voice, sound effects, ambience and music modeled together
Describe sounds and dialogue in your prompt and the model generates them natively
๐จ NSFW Support
Comprehensive content generation without restrictions
Dedicated NSFW presets tuned for LTX 2.3 I2V and FL2V
Natural progression, consistent characters, explicit diegetic soundscape
๐ฆ What's Included โ 2 Workflows
๐ผ๏ธ I2VA โ
LTX23-I2VA-Qwen3.5.jsonโ text + first-frame image (image) โ Image-to-video+audio. 10Eros v1.5 + DMD hybrid v2 LoRA, dual-stage with spatial upscaler.๐ FL2VA โ
LTX23-FL2VA-Qwen3.5.jsonโ text + first-frame (image) + last-frame (image2) โ First-Last-Frame to video with audio. Qwen3-VL sees both frames and describes the transition. Includes TensorRT upscale + RIFE frame interpolation for high-resolution output.
Both workflows include TensorRT upscaling (RealESRGAN x4) and RIFE frame interpolation.
๐ผ๏ธ Multi-Reference Input (image2)
The QwenVL-Mod node has two image inputs:
I2VA:
image= first frameFL2VA:
image= first frame,image2= last frame,frame_count = 1
Qwen3-VL sees all connected images as individual images (not as a video sequence), enabling proper multi-reference analysis for FL2VA.
๐ฏ QwenVL-Mod NSFW Presets
The workflows include built-in NSFW presets for the Qwen3-VL prompt enhancer:
๐ฌ LTX 2.3 NSFW I2Vโ Image-to-video with mandatory audio instructions๐ LTX 2.3 NSFW FL2Vโ First-last-frame transition with audio instructions
What the presets produce
๐ฌ I2VA: Rich scene description, camera motion, subject action, explicit diegetic soundscape, non-diegetic music flag
๐ FL2VA: Describes the transition path between first and last frames (not the scene โ images fix the scene). Favors single continuous shot.
All presets: smooth, continuous camera motion (no abrupt or stepped changes), explicit audio notes
SFW presets are also available. Edit the preset dropdown in the QwenVL node to switch.
๐ฎ Usage Examples
Image-to-Video (I2VA)
Load
LTX23-I2VA-Qwen3.5.jsonUpload your first-frame image to
imageSelect preset
๐ฌ LTX 2.3 NSFW I2VWrite what happens next (in any language)
Generate animated video with synchronized audio
First-Last-Frame (FL2VA)
Load
LTX23-FL2VA-Qwen3.5.jsonUpload first-frame to
image, last-frame toimage2, setframe_count=1Select preset
๐ LTX 2.3 NSFW FL2VDescribe the transition between the two frames
Generate the interpolated video with TensorRT upscale + RIFE
๐ง Technical Specifications
โก Performance
Output: Native LTX 2.3 resolution, 24 fps
Audio: Native synchronized stereo audio
Upscale: TensorRT RealESRGAN x4 (both workflows)
Frame interpolation: RIFE rife49 โ higher fps (both workflows)
Sage Attention: FP16 accumulation, async offload
Smart caching: Reuse prompts with same inputs, Fixed Seed Mode for text-only caching
๐จ Model Support
Qwen3-VL 4B: 7 GGUF variants (2.38 GB โ 4.28 GB)
Qwen3-VL 8B: 7 GGUF variants (4.8 GB โ 8.71 GB)
Qwen3.5: 4B / 9B / 27B (uncensored, heretic, unsloth) โ thinking mode disabled
HF Models: Josiefed, official, Heretic-Stable variants
Quantization: Q4_K_S, Q5_K_S, FP16, INT8
๐ Multilingual Capabilities
Input languages: Any language supported
Auto-translation: Automatic translation to optimized English
Style detection: Works with multilingual prompts
Cultural adaptation: Context-aware prompt enhancement
๐ฆ Installation
Quick Install
Download: ComfyUI-QwenVL-Mod (latest version)
Extract to
ComfyUI/custom_nodes/ComfyUI-QwenVL-ModInstall requirements:
pip install -r requirements.txtRestart ComfyUI
Load included workflows from
ltx/23/folder
Custom Nodes Required
ComfyUI-QwenVL-Mod โ All workflows (Qwen3-VL prompt enhancer) โ huchukato/ComfyUI-QwenVL-Mod
ComfyUI-LTXVideo โ LTX 2.3 native nodes โ Lightricks/ComfyUI-LTXVideo
ComfyUI-RIFE-TensorRT-Auto โ FL2VA, I2VA (frame interpolation) โ huchukato/ComfyUI-RIFE-TensorRT-Auto
ComfyUI-Upscaler-TensorRT-Auto โ FL2VA, I2VA (upscaling) โ huchukato/ComfyUI-Upscaler-TensorRT-Auto
ComfyUI-VideoHelperSuite โ (VHS_VideoCombine) โ Kosinkadink/ComfyUI-VideoHelperSuite
ComfyUI-Easy-Use โ (easy showAnything) โ yolain/ComfyUI-Easy-Use
comfyui-find-perfect-resolution โ All workflows (ResolutionSelector) โ ashtar1984/comfyui-find-perfect-resolution
was-node-suite-comfyui โ (ComfyMathExpression) โ ltdrdata/was-node-suite-comfyui
Models Required
LTX 2.3 I2VA / FL2VA use the 10Eros v1.5 uncensored setup:
models/checkpoints/โ10Eros_v1.5_fp8mixed_experimental_learned.safetensors(~30 GB) โ LokkenJP/10EROS_1.5_fp8_exp_learnedmodels/text_encoders/โgemma-3-12b-it-ablit-norms-biproj-fp8mixed.safetensors(~12.8 GB) โ TenStrip/LTX2.3-10Erosmodels/loras/ltx23/โLTX2.3_DMD_hybrid_v2.safetensors(~662 MB) โ TenStrip/LTX2.3_DMD_Loramodels/latent_upscale_models/โltx-2.3-spatial-upscaler-x2-1.1.safetensors(~996 MB) โ Lightricks/LTX-2.3models/latent_upscale_models/โltx-2.3-temporal-upscaler-x2-1.0.safetensors(~262 MB) โ Lightricks/LTX-2.3models/vae/โpruna_ltx2.3_vae_comfy_bf16.safetensors(~500 MB) โ Kijai/LTX2.3_comfy
Qwen3-VL Prompt Enhancer
models/LLM/โQwen3-VL-8B-Heretic-Stable(GGUF or HF)
TensorRT Engines
models/upscale_models/โRealESRGAN_x4(TensorRT engine)models/rife/โrife49_ensemble_True_scale_1_sim(TensorRT engine)
TensorRT engines must be built for your specific GPU. See ComfyUI-RIFE-TensorRT-Auto and ComfyUI-Upscaler-TensorRT-Auto for build instructions.
๐ณ Docker / Cloud Ready
OneClick RunPod Template
Prefer a ready-to-go environment? Use the OneClick - ComfyUI - LTX 2.3 Uncensored - CU13 RunPod template:
Docker image:
huchukato/comfyui-qwenvl-runpod:cu13-ltx23Base:
runpod/comfyui:cuda13.0All custom nodes pre-installed
Both workflows auto-downloaded at boot
Models auto-downloaded at first boot (~40 GB, persistent)
ComfyUI v0.32.0+ forced at boot
Sage Attention, FP16 accumulation, async offload
TensorRT upscaling + RIFE frame interpolation
ComfyUI Args (pre-configured)
--disable-auto-launch
--fast fp16_accumulation
--use-sage-attention
--cuda-malloc
--async-offload
๐ Why Choose ComfyUI-QwenVL-Mod + LTX 2.3?
๐ฌ For Content Creators
Native audio: Video and audio in one pass โ no separate audio node
Multilingual: Write in any language, Qwen3-VL handles translation
Professional: Smooth camera vocabulary and natural motion
Quality: 10Eros v1.5 uncensored + TensorRT upscale
๐ฅ For NSFW Content
Explicit: Uncensored generation with dedicated NSFW presets
Detailed: Rich scene descriptions with explicit diegetic soundscape
Natural: Realistic progression, consistent characters
Audio: Native moans, breaths, skin contact, ambient sounds
โก For Power Users
Customizable: Easy to modify presets and system prompts
Extendable: Add your own Qwen3-VL models (GGUF or HF)
Integrable: Works with existing ComfyUI setups
Optimized: Sage Attention, FP16, async offload, smart caching
Multi-reference:
image2input for FL2VA workflow
๐ What Makes This Special?
First: Complete LTX 2.3 uncensored workflow pack with Qwen3-VL auto-prompting
Native audio: No separate audio node โ LTX 2.3 does it all
2 workflows: I2VA + FL2VA covering the main LTX 2.3 use cases
Multi-reference: Qwen3-VL sees all connected images (not just the first)
TensorRT: Built-in upscaling and frame interpolation
NSFW presets: Dedicated presets tuned for LTX 2.3
Multilingual: Any input language, auto-translated and formatted
Ready: Works out-of-the-box with included workflows
๐ฏ What's New in v2.3.1
โ LTX 2.3 10Eros Setup
Switched to 10Eros v1.5 uncensored checkpoint
Added TenStrip FP8 abliterated text encoder (
gemma-3-12b-it-ablit-norms-biproj-fp8mixed.safetensors)Removed obsolete Gemma 3 FP4 + abliterated LoRA combo
Added DMD hybrid v2 LoRA for fast sampling
โ Native Audio Workflows
Both I2VA and FL2VA workflows generate synchronized native audio
Presets include mandatory audio instructions
โ Qwen3.5 Rename
Workflow filenames updated from
Qwen3VLtoQwen3.5for consistency