Yet Another Workflow : easy t2v + i2v
I've aimed at a user-friendly UI for ComfyUI. There's a balance between complexity and ease of use, and this workflow aims to give you useful controls with clear guidance on what you need to care about. I hope these will be helpful to anyone strugging with quality and the general UI-isms of ComfyUI. I've taken the time to color code and add lots of notes. Please read the notes, I've tried to make them useful!
This is the workflow I use, it's not aimed at a skill level. It's designed to be easy to use and adjust with some UI concessions and labeling to ensure you can pilot it with less experience in a way that is more sophisticated than the official example workflows, which can be easy to break.
The primary goal with this workflow is to give you a strong foundational place to generate either text to video (T2V) or image to video (I2V) outputs without having to fuss too much.
The green controls are the stuff you generally want to mess with.
The secondary goal here is to provide a consistent interface to interact with different samplers and now different models.
Now for LTX-2.3
One of the strengths of respecting a consistency to a UI is that it allows you to change things under the hood to enable people to explore new techniques without changing the majority of their experience. Whether you're familiar with my Wan 2.2 workflow or whether this is your first experience with my workflow, once you're comfortable with one of them, the goal is that you can switch between them with minimal fuss.
LTX-2.3 currently represents the future of open weights video. Things are always in flux (pun intended), but until Wan announces a new open weights release, LTX-2.3 is the actively developing model. It represents both steps forward and backwards. It introduces a number of long desired features:
Native support for sound (both prompted and from files)
Long video generation
Better support for lower end hardware (kinda)
Single model process presents a visually simpler process for I2V and T2V
As before there are lots of notes and color coding to help your orient yourself. If you're new to this workflow, but not to YAW in general, I've added a bunch of usage notes that specificaly relate to LTX-2.3.
The LTX-2.3 release is notable becaue it made some major improvements to audio, prompt adherance, and LoRA training.
I have a lot of criticisms of LTX-2.3, and the RunPod guide will cover more of the details, but my major complaint is prompt adherance and verbosity. It wants very verbose prompts and is bad at infering what you want. On top of that, the way they achieve prerformance improvements makes prompt adherance markedly poorer. You'll end up doing way more gens with LTX-2.3 than something like Wan 2.2. Recently the contributions of TenStrip and Alissonerdx have made some material improvements.
What's new?
This is probably the longest development iteration. The situation with LTX-2.3 has been shifting for a while. I have done extensive testing across a broad set of workflows and techniques. I've had three different version that we close to release candidates, but then something would change! But now we are here! This new setup gives you a bunch of new options to control and tweak your generations while providing you with a reliable baseline to work from.
This new workflow moves away from the official LTX-2.3 methodology. Special thanks here to @tenstrip for their relentless innovation in LTX-2.3. Their work on the DMD LoRA which merges the official 384 1.1 distill LoRA with the data from JoyAI's Echo really improves prompt adherance. This workflow hews closely to his generation strategy. Additional gratitude to Alissonerdx for the Best FaceID LoRA that improves I2V consistency.
In addition to migrating the adjustments made during the Wan 2.2 update to v0.50, there are some minor unique additions to the LTX workflow which include the option to include a watermark (which is a speaker icon by default). Many notes have been adjusted to better describe how LTX-2.3 works.
Like it?
Give it a like! Tag it as a Resource when you use it! Support on Patreon or a tip on Ko-fi are also welcome. Yellow Buzz will go towards promoting awareness here on Civit.
Need help?
I like helping people get going with this stuff, so if you want help message me. If you want extended one-on-one help, there's an option on the Patreon. I'm happy to walk you through the details, answer your questions, and give you some extra tips and tricks, and scripts. I've done this for a few folks, I'll save you money and headaches.
I've also written an article here on getting it going with my Runpod template. The template will vastly expedite and simplify getting things up and running.
General Advice
Make lots of videos! Post your videos! Don't fuss with the tech! Be smart about how you spend your time with this stuff. It's easy to burn out if you spend more time trying to get things to work than making videos you like. That's really why I'm posting this.
Use RunPod. Use the RTX 5090 or PRO 6000 or the H100 SXM. Use my LTX-2.3 template. If you've not used RunPod before, sign up with my link; we'll both get some free credit. See the article for more.
If you use a service like RunPod, if you're doing I2V, it can be smart to have your images ready in advance to make sure the server stays busy while you are using it.
If you run this outside of Runpod, you'll need to install some custom nodes. To do that, click the "Manager" button at the top of the Comfy interface, and then click the "Install Missing Custom Nodes". Click "Install" on each one - I recommend in order; you'll need to wait till each has installed. Do not bother restarting ComfyUI until they are all installed. The RunPod template has them preinstalled.
If the wires bother you, there's a button in the bottom right on the floating UI that will hide them.
This workflow is setup for .safetensors models, but you can use GGUF if you want to make the changes node changes.
Costs?

Here's my most recent benchmark evaluation for LTX-2.3. It has a fairly linear peformance profile, which is to say: you mostly get what you pay for. As noted, the RTX 5090 is a decent deal, but the PRO 6000 works out to a similar cost. You'll have a higher spend, but much faster video gen. H100 is a good performance choice. Quite nimble, if you want to make videos quickly.
Troubleshooting
If a node is missing (bright thick red outline with a warning when you open the workflow), you can install them by going to Manager > Install Missing Custom Nodes, and pressing Install on any the nodes that show up there.
If you are getting any errors related to a custom node, it's possible something has changed recently in the software. It might be useful to change a version back to the last "stable" build in these situations.
For example, the nightly build of WanVideoWrapper might introduce an error that wasn't there last time. With a workflow open, you can go to Manager > Custom Nodes in Workflow. This will show you all of the custom nodes. If you click, Switch Ver, you can see all of the releases. Consider trying the first numbered on at the top of the list.
If that doesn't work, or there seem to be more significant problems and you are using RunPod, you may have forgotten to select CUDA 12.8. Try restarting the server. If that doesn't work, terminate the pod, and make a new one. This will fix a surprising number of possible issues.
Longer video generation support?
LTX-2.3 supports long video generation (memory limits apply). The longer the video, the more wonky the prompt adherance becomes (more unexpected physics, garbled dialog, etc). I've set a maximum on the length slider, but you can increase it if you want to go nuts. I've done over a minute so far. The results are not great, but with enough memory you can do it.
Sound
LTX-2.3 has decent sound support, but it's a double edged sword. Sometimes the motion will be great and the sound gets weird. Sometimes the sound is great, but the character does something strange or an extra hand appears or some such. There's no good way to fix this. I want to call out that there's not a great way way to get a consistent voice across prompts.
Description
The intial release of YAW for LTX-2.3:
Initial adaptation for LTX 2.3 including the full sweet of official sampler options and changes to facilitate their different nodes
Created a new image size custom node to reduce size adjustments for i2v
Switching node isolate to improve install issues
FAQ
Comments (15)
I downloaded the workflow and I like it. It runs smooth without any issues.
Even just using the 1st Sampler alone produces some cool videos.
Thanks for the kind feedback. Yeah, that's the intent. I think I might need to add some extra signaling. The note explains this, but it's not SUPER obvious that just doing the first sampler will give you nice looking videos.
what is it with ltx that it always wants to zoom in really close and out again. I have that in so many videos. can´t even prompt against it
I would double check you are following the LTX prompting style recommendations. You should be able to prompt a static camera. But as I've noted in a few places, LTX is tricky with prompt adherance, so ever more important to make sure you're doing what is expected when you encounter something you don't want. (Tangentially related, there are some official camera LoRA's that can force specific camera movements if you need them. )
out of the 20 tries all had that zoom in. despite even prompting against it. litke static camera or no camera movement. Its almost as if it is overtrained on zoom ins.
@chrisbraeuer41172035 Mmm. I'll have to check mine and see how consistent it is with mine.
Prompt adherence...sigh.
It really is a struggle trying to get what you want. In comfyui preview nodes you can see as its generating that it alters your Gen if it does not understand or like agree what youre asking. Its has to be built in mechanisms within the transformer that are fighting nsfw prompts ESPECIALLY T2V!!!
Cant get anything out of that!
I don't think that bit about the NSFW is true. It's just bad at prompt adherance in general. It's part of their tech stack decision to ensure it runs on cheap cards. Not sure if that was worth it but here we are. Even with an abliterated gemma model, it still struggles to follow fairly clear requests - NSFW or not. It will often do what you want eventually, but it can take many attempts.
I have a lot of experience with your WAN 2.2 workflow and it’s great.
I just tried a LTX 2.3 workflow from one of the public runpod templates and it worked better than I expected so I’m thinking about trying your LTX 2.3 workflow now.
Is 10Eros kind of the equivalent of Smooth Mix in that it bakes in some of the LTX loras? or is there a better analogy for thinking about it? If I add that checkpoint to CHECKPOINT_IDS_TO_DOWNLOAD, is it essentially a drop-in replacement in an existing node in the workflow, or would it involve adjustments to the workflow itself?
Thanks!
Kind of! 10eros is an adjustment of the Sulphur data to optmize for I2V. (Sulphur is available as both a checkpoint - a mix of a base model with LoRA data - and as a LoRA.) That data is custom, not merely a merge of a bunch of public LoRA's. Smooth Mix is a combination of a bit of custom data with public data. Where as something like DiSiWa is essentially just a merge of public data.
In general I don't recommend checkpoints, but in this case it's fine. It's all custom data, but you do lose the ability to adjust the strength. They both (Sulphur and 10eros) mix nicely with other LoRAs.
The checkpoint is a replacement for the base Dev model in this case. You'd just select it in the places that refer to the checkpoint in the model loading sections. (No nodes are replaced, just adjust the settings.)
I'm using your YAW LTX-2.3 and I could not get it going. It did not install models on startup, the excellent "runpoddownloader" did not work (it works perfectly on YAW Wan 2.2). I'm new to this but I just could not get a handle on it.
Do you happen to have a log file? Unless you set all of the models to false or tried to download from CivitAI without a key, it should work. It's not doing the same process as the Wan template yet. (Will come sooner than later.)
can you all fucking decide between checkpoint, unet and diffusion models folder .... for fucks sake
you know what, your workflow is retarded. It's pulling audio VAE from checkpoints....
We're not one person. Different people have different needs. I've made my choices for my reasons, other people make different choices.
The original workflow was built for the base model, not a checkpoint. It follows the way the official workflows work. The next version, which will release soon, will work slightly differently.
Also there's absolutely nothing wrong with using the audio VAE of the checkpoint.