━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
✨ Krea2 → LTX-2.5 Motion Suite — Text Prompt to Video with Native Synced Audio (SFW · 16GB)
ComfyUI · Krea-2 Community License + LTX-2.x Community License · Turbo T2I → LTX-2.5 I2V with native audio
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Type a SUBJECT and a POSE, press Queue, get a still image and an animated video with native synced audio — no separate tools, no manual image export/import between workflows. This graph fuses Krea2 Turbo (fast DiT text-to-image) with LTX-2.5 (image-to-video with native synchronized audio) in one seamless funnel: Krea2 renders the starting frame, a VRAM-unload bridge clears it from memory, built-in QwenVL prompt enhancement turns your short SUBJECT + POSE into a photographic prompt (and separately, an automatic motion caption from the rendered frame), then LTX-2.5's distilled two-stage pipeline animates it with native audio decoded in the same diffusion pass — no bolt-on sound generation. This workflow is SFW-only by design, tuned for clean scenic/portrait content that complies with Gemma's Prohibited Use Policy (Gemma-4 powers the LTX-2.5 text encoder).
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
✅ SFW-Only Disclosure
This workflow's QwenVL prompt-enhancer system and output generation are tuned exclusively for SFW scenic, portrait, and nature content. No mature/NSFW keywords in the enhancer templates; no adult-oriented showcase examples. All example showcase videos depict AI-generated clean scenes (outdoor locations, nature, light motion) without explicit content. This listing and all outputs are fully compliant with Gemma-4's Terms of Use and Prohibited Use Policy (which restricts sexually explicit content, dangerous/violent content, hate speech, etc.). Users are responsible for following Civitai's Terms of Service and their local jurisdiction's laws when publishing generated content.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
✨ Features
✅ One Graph, Zero Handoff — SUBJECT + POSE prompt in, still image AND animated video with audio out. Krea2 generates the frame internally and hands it straight to LTX-2.5 — no LoadImage node juggling between separate workflows.
✅ SUBJECT + POSE Boxes with Built-In Examples — two clearly labeled, resizable text boxes (SUBJECT, POSE/ACTION) sit at the top of the graph, each with an example-list Note right beside it — copy an example in, or write your own. No hunting through the graph for what's tunable.
✅ VRAM-Safe Double Bridge — easy cleanGpuUsed unloads Krea2 before QwenVL's motion-caption pass loads, then unloads QwenVL again before LTX-2.5 loads — the three models never co-reside in VRAM.
✅ Native Synced Audio — LTX-2.5's joint A/V latent space decodes video and audio from the same diffusion pass. The audio output is synchronized to the motion without a separate audio model — footsteps, water, wind, fire, whatever the scene needs, all matched to the motion.
✅ Two-Stage Distilled Pipeline — Stage 1: 8-step draft at half-res. Stage 2: 2x spatial upscale + 3-step refine to full res. Distilled model keeps this runnable on 16GB VRAM with no separate LoRA needed.
✅ Auto-Orientation — reads your rendered source image and picks portrait/landscape automatically, no manual toggle.
✅ Dual Motion-Caption Mode — Motion Switch: 0 = QwenVL auto-captions the scene straight from your frame (visual + audio description), 1 = type your own motion+audio text manually for tighter control.
✅ Dual Orientation — one switch sets both stages together: 1920×1088 landscape (16:9) or 1088×1920 portrait (9:16) — Krea2 and LTX-2.5 always stay in matching aspect.
✅ Krea-2 Original Architecture — NOT FLUX-derived; DiT 12.9B model from Krea, FP8 quant by AlperKTS.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
📦 Required Models (8 files, ≈51 GB total)
Krea2 T2I stage (~10 GB):
• krea2_turbo_fp8.safetensors (~7 GB) — Krea2 Turbo DiT 12.9B UNet, FP8 quant by AlperKTS
• qwen3vl_4b_fp8_scaled.safetensors (~2–3 GB) — Krea2 text + vision encoder (Qwen3-VL-4B, FP8)
• qwen_image_vae.safetensors (~1 GB) — Krea2 image VAE
LTX-2.5 I2V stage (~40 GB, int8-convrot variant):
• ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors (~21.5 GB) — LTX-2.5 DiT, distilled 22B, int8-convrot quant
• gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors (~15.4 GB) — LTX-2.5 text encoder (Gemma-4, int8-convrot)
• ltx-2.5-video-vae-bf16.safetensors (~1.4 GB) — LTX-2.5 video VAE
• ltx-2.5-audio-vae-bf16.safetensors (~0.4 GB) — LTX-2.5 audio VAE (native audio decode)
• ltx-2.3-spatial-upscaler-x2-1.1.safetensors (~1.6 GB) — stage1→stage2 latent upscaler bridge
Prompt Enhancement (auto-downloaded on first use):
• Qwen3-VL-2B-Instruct (~2.5 GB, auto-cached) — VLM enhancer for SUBJECT/POSE→prompt and frame→motion caption; downloaded automatically by ComfyUI-QwenVL on first run, not manually placed
Total disk usage: ~51 GB models + ~2–5 GB working files (VAE cache, temp outputs). Qwen3-VL-2B auto-downloads on first run (~2.5 GB).
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
⬇️ Download Links
📁 ComfyUI/models/unet/Krea2/
• krea2_turbo_fp8.safetensors — https://huggingface.co/AlperKTS/Krea2_FP8
📁 ComfyUI/models/clip/ (or text_encoders/)
• qwen3vl_4b_fp8_scaled.safetensors — https://huggingface.co/AlperKTS/Krea2_FP8
📁 ComfyUI/models/vae/
• qwen_image_vae.safetensors — https://huggingface.co/AlperKTS/Krea2_FP8
• ltx-2.5-video-vae-bf16.safetensors — https://huggingface.co/Lightricks/LTX-2.5/blob/main/vae/ltx-2.5-video-vae-bf16.safetensors
• ltx-2.5-audio-vae-bf16.safetensors — https://huggingface.co/Lightricks/LTX-2.5/blob/main/vae/ltx-2.5-audio-vae-bf16.safetensors
📁 ComfyUI/models/diffusion_models/
• ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors — https://huggingface.co/Lightricks/LTX-2.5/blob/main/diffusion_models/ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensors
📁 ComfyUI/models/text_encoders/
• gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors — https://huggingface.co/Lightricks/LTX-2.5/blob/main/text_encoders/gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors
📁 ComfyUI/models/latent_upscale_models/
• ltx-2.3-spatial-upscaler-x2-1.1.safetensors — https://huggingface.co/Lightricks/LTX-2.3/blob/main/ltx-2.3-spatial-upscaler-x2-1.1.safetensors
⚠️ Filenames above are exactly what the workflow JSON expects in its loader nodes — match them exactly, or update the loader node if your local copy is named differently.
Auto-Downloaded Models (no manual placement needed):
• Qwen3-VL-2B-Instruct — automatically downloaded by ComfyUI-QwenVL custom node on first workflow run (~2.5 GB, cached afterward).
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
🧩 Required Custom Nodes (3 packs)
1. ComfyUI-QwenVL (1038lab / AILab) — provides AILab_QwenVL_PromptEnhancer (SUBJECT+POSE → photographic prompt) and AILab_QwenVL (frame → motion caption / audio description). License: Apache-2.0/BSD. https://github.com/1038lab/ComfyUI-QwenVL
2. ComfyUI-Easy-Use (vjumpkung fork) — provides easy cleanGpuUsed (VRAM-unload bridge) and easy anythingIndexSwitch (Motion/Orientation switches). https://github.com/vjumpkung/ComfyUI-Easy-Use
3. ComfyUI-WAS-Node-Suite (WASasquatch) — provides Text Multiline (the SUBJECT and POSE input boxes). License: MIT. https://github.com/WASasquatch/was-node-suite-comfyui
LTX-2.5's own nodes (LTXVConditioning, LTXVPreprocess, LTXVConcatAVLatent, LTXVSeparateAVLatent, LTXVAudioVAEDecode, etc.) are native to ComfyUI ≥0.30 — no extra custom node pack needed for those, just an up-to-date ComfyUI install.
Install via ComfyUI Manager (search each pack name, or use "Install Missing Custom Nodes" after loading the JSON).
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
🚀 How to Use
Quick Start:
1. Download all 8 manually-placed model files → place in ComfyUI/models/ (see paths above)
2. Install the 3 custom node packs via ComfyUI Manager
3. Confirm ComfyUI ≥0.30
4. Load the workflow JSON into ComfyUI
5. Type your scene into the SUBJECT box (top of graph) — copy one of the example lines from the Note beside it, or write your own
6. Type your pose/action into the POSE / ACTION box beside it — same idea, examples right there
7. Leave Motion Switch = 0 (auto-caption) for the simplest path
8. Queue → Krea2 renders the still → cleanGpuUsed unloads it → QwenVL auto-captions motion/audio from the frame → cleanGpuUsed unloads QwenVL → LTX-2.5 stage1 (8-step draft) → stage2 (2x upscale + 3-step refine) → dual VAE decode (video + audio) → MP4 lands in your output folder
What's happening under the hood:
- Krea2 T2I: SUBJECT + POSE (combined, expanded by QwenVL PromptEnhancer) → EmptySD3LatentImage (1920×1088 or 1088×1920) → KSampler (euler, simple, 10 steps, cfg 1.0) → VAEDecode
- Bridge #1: easy cleanGpuUsed clears Krea2 from VRAM
- Motion caption: QwenVL auto-captions rendered frame into motion/atmosphere + audio description (or your own Manual Motion text, via switch)
- Bridge #2: easy cleanGpuUsed clears QwenVL from VRAM before LTX-2.5 loads
- LTX-2.5 Stage1 (draft): ImageScale to 960×544 or 544×960 → LTXVConditioning → LTXVPreprocess → SamplerCustomAdvanced (euler_ancestral, 8 steps, CFG 1.0) → LTXVLatentUpsampler (2x spatial bridge to stage2)
- LTX-2.5 Stage2 (refine): 2x-area scaled latent → LTXVConditioning → LTXVPreprocess → SamplerCustomAdvanced (3 steps, same sampler) → LTXVSeparateAVLatent (split video + audio)
- Output: VAEDecode (video, bf16) + VAEDecodeAudio (audio, bf16) → CreateVideo (24fps, stereo audio) → SaveVideo → H.264 MP4 with synced audio track
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
⚙️ Settings & Parameters
• Krea2 Sampler — euler, simple scheduler, 10 steps, CFG 1.0
• Krea2 Resolution — 1920×1088 (16:9) or 1088×1920 (9:16), via Orientation Switch
• LTX-2.5 Stage1 — 960×544 or 544×960 (half-res draft), euler_ancestral, 8 steps, CFG 1.0
• LTX-2.5 Stage2 — 1920×1088 or 1088×1920 (full-res after 2x upscale), 3 steps, same sampler
• Frames — 241 frames @ 24fps (~10.0s) — valid frame counts: 49/73/97/121/145/169/193/217/241 (N×8+1 constraint from LTX-2.5)
• Motion Switch — 0 = Auto (QwenVL captions the frame) / 1 = Manual (type your own MOTION + AUDIO text)
• Orientation Switch — 0 = Landscape / 1 = Portrait (sets both Krea2 and LTX-2.5 automatically)
• Negative Prompt — curated quality filter, includes anatomy guards (bad anatomy, extra limbs, malformed hands)
• VRAM Bridges — easy cleanGpuUsed ×2 (do not remove — these keep Krea2 / QwenVL / LTX-2.5 from co-residing in VRAM)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
✍️ Writing SUBJECT + POSE (SFW Scenic Recipe)
Two boxes, two jobs: SUBJECT = what + where; POSE / ACTION = how it moves/unfolds. QwenVL combines and expands both into one scene description automatically.
SUBJECT examples (copy one in, or use as a template):
• vibrant garden in spring bloom, winding stone path
• rocky coastline at sunset, waves crashing against cliffs
• golden wheat field under afternoon light
• misty forest with ancient trees, soft morning light
• mountain lake reflecting alpine peaks
• wildflower meadow in motion, swaying in breeze
• desert landscape at dusk, warm amber light
• forest stream with smooth water flow, moss-covered rocks
POSE / ACTION examples:
• gentle pan across the scene, slow fluid motion
• camera slowly tracking along the path
• wide establishing shot, subtle floating motion
• close-up of water details with light reflection
• sweeping view left to right over landscape
• pull-back from foreground to reveal full vista
• time-lapse style progression of light change
• birds-eye view descending into a valley
Mix and match freely — SUBJECT and POSE are independent boxes, so any combination works. Short inputs are fine; QwenVL fills in lighting, weather, and motion detail automatically. For best audio sync, choose scenes with natural sound (water, wind, leaves, weather). For guaranteed specific sounds, use Manual Motion mode (Motion Switch=1) and write your own "VISUAL... AUDIO: [sound list]" text.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
🎨 Showcase
Audio-synced hero themes — real sound transients timed to on-screen motion, not ambient filler. 1920×1088, 24fps, native audio kept (not muted), base checkpoint only (no extra LoRAs), motion auto-captioned by QwenVL. All 6 clips: 10.04s video + 10.01s audio (this file's default length).
1. 🎻 Street Violinist — a violinist plays under warm city night lights, bow arm sweeping across the strings. Audio: bow-stroke transients sync to the arm motion, not looped ambient street noise.
2. 🐎 Horse Gallop on Beach — hooves splash through shallow surf at golden hour, mane flowing, no rider. Audio: hoof-splash transients sync to footfall timing.
3. 🌧️ Neon Rain Walk — a coated pedestrian walks a rain-slicked night street, neon signage reflected in the puddles underfoot. Audio: rain + puddle-splash transients sync to each footfall.
4. 🔥 Campfire Under the Stars — a camper keeps watch beside a crackling campfire in a forest clearing, embers rising into the night sky. Flames and sparks animate genuinely across the full 10s take.
5. 💃 Golden-Hour Dancer — flowing silk dress catches the light mid-motion at golden hour, natural relaxed movement.
6. 🍁 Koi Pond Zen Garden — a quiet moment beside a traditional koi pond at dawn, fish gliding beneath the surface, maple leaves drifting down through soft mist.
⚠️ Known limitation — dynamic point-light VFX: an earlier ⛈️ thunderstorm/lightning theme was dropped from this showcase. LTX-2.5 rendered the bolt as a static light streak baked into the source still — it did not animate (no flash/flicker) across the clip, unlike the fire and water motion above. If your SUBJECT/POSE text calls for lightning, fireworks, or similar bright instantaneous flashes, expect the same static-glow limitation and plan to iterate seed/prompt or pick a different motion type.
All videos generated with the full funnel (Krea2 T2I → LTX-2.5 I2V) using the base checkpoint only (no extra LoRAs). Motion auto-captions via QwenVL unless noted otherwise.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
💡 Performance Tips
• Minimum Spec — 16 GB VRAM and ~32 GB system RAM. The double cleanGpuUsed bridge offloads models to system RAM between stages; low-RAM systems may thrash during handoff.
• VRAM — Tested on RTX 5080 16GB — peak ~15.2 GB across real batch runs (int8-convrot, native resolution). Treat 16GB as a hard minimum, not a comfort margin.
• Cold Start — On first workflow run, Qwen3-VL-2B-Instruct (~2.5 GB) auto-downloads and caches. Plan an extra 30–60 seconds the first time only.
• Restart ComfyUI Before First Full Run — RAM accumulates across separate stage testing. Start with a clean ComfyUI session before your first production run of the complete funnel.
• SUBJECT / POSE Split — keep scene/setting in SUBJECT and camera/motion in POSE; blending them into one box tends to confuse the enhancer.
• Motion Auto-Caption Quality — QwenVL reads the rendered still, so a clearer, well-lit SUBJECT composition tends to produce richer audio-tailored captions.
• SFW Compliance — this workflow is tuned for clean scenic content. Steer SUBJECT toward nature, outdoor, portrait, and architecture prompts for best results and fullest Gemma Terms compliance.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
📝 Notes & AI Disclosure
• AI-Generated Content — All example outputs are AI-generated by Krea2 Turbo + LTX-2.5. Respect local AI disclosure laws when publishing your own outputs.
• Hardware — Tested on RTX 5080 16GB — peak VRAM ~15.2GB across real batch runs (see Performance Tips above).
• Configuration Only — no model weights embedded in the JSON; download all 8 manually-placed files separately from the sources listed above.
• Workflow Reuse — feel free to modify, share, and fork.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
🔗 Also Check Out
🔗 Sister workflow: [Krea2 → MiniMax H3 Motion Suite](https://civarchive.com/models/2838477) — the NSFW/H3 video variant of this funnel if you want mature-content support.
🔗 Sister workflow: [LTX-2.5 Image-to-Video AUTO](https://civarchive.com/models/2853322) — the standalone i2v-only stage if you already have reference images and want to skip the T2I pass.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
## Changelog
v1.0 (2026-08-15)
Initial release — funnel merge of Krea2 Turbo T2I (base checkpoint, FP8 quant) + LTX-2.5 I2V (22B distilled, int8-convrot quant, two-stage pipeline) with native synced audio. SFW-only design, QwenVL auto-prompt + auto-motion-caption, double VRAM-safe bridge, auto-orientation. Tested on RTX 5080 16GB, peak ~15.2GB. Includes the temporal_size=128 micro-stutter fix from LTX-2.5 v1.4 and the "Cockpit-Left / Pipeline-Right" layout standard.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
⭐ Found this useful?
• Like if it saved you time turning scripts into videos with real sound
• Comment your results — I read every one
• Follow for new ComfyUI workflows, all tested on 16 GB VRAM
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
⚖️ Model Attribution & Licensing
Krea-2 Turbo (AlperKTS FP8 Quant)
• License: Krea-2 Community License — https://www.krea.ai/krea-2-licensing
• Commercial use: OK if total annual revenue < $1,000,000 USD
• Model: original architecture, DiT 12.9B, not FLUX-derived
LTX-2.5 (Lightricks, distilled 22B, int8-convrot quant)
• License: LTX-2.x Community License Agreement — https://github.com/Lightricks/LTX-2/blob/main/LICENSE.md
• Commercial use: Free for entities under $10M annual revenue; paid Commercial Use Agreement required at ≥$10M ARR
• Territory: No explicit territory restriction beyond standard US export/sanctions compliance
• Content: This workflow is SFW-only by design; outputs fully compliant with Gemma Prohibited Use Policy
Gemma-4 Text Encoder (Google)
• License: Gemma Terms of Use (https://ai.google.dev/gemma/terms) and Prohibited Use Policy (https://ai.google.dev/gemma/prohibited_use_policy)
• Scope: The Gemma Terms of Use cascade downstream to anyone using this workflow. By using this workflow you agree to Gemma's Terms of Use. This workflow's SFW scenic content is fully compliant with the Prohibited Use Policy (which restricts sexually explicit content, dangerous/violent content, hate speech, etc.).
• Verified: 2026-08-15
Supporting Components — all Apache-2.0
• Qwen3-VL-4B (Krea2 text/vision encoder) • Qwen Image VAE • Qwen3-VL-2B-Instruct (prompt enhancer model) • LTX-2.5 video VAE • LTX-2.5 audio VAE
ComfyUI Custom Nodes
• ComfyUI-QwenVL (1038lab) — Apache-2.0/BSD
• ComfyUI-Easy-Use (vjumpkung fork) — per upstream repository
• ComfyUI-WAS-Node-Suite (WASasquatch) — MIT
Workflow JSON — original work, free to use, modify, and redistribute.
Full attribution detail (sources, license text, verification dates) is maintained in this listing. All example outputs are AI-generated. Model weights remain the property of their respective owners; download separately from the official sources above.
Description
Initial release — SUBJECT + POSE prompt in, animated video with native audio out. Krea2 Turbo T2I feeds LTX-2.5 I2V (int8-convrot, two-stage distilled) via a double VRAM-safe cleanGpuUsed bridge, with QwenVL auto-prompt + auto-motion-caption built in. SFW-only, 16GB VRAM target. Includes temporal_size=128 micro-stutter fix and "Cockpit-Left / Pipeline-Right" layout standard.