These models are the result of a partial NVFP4 quantization of Wan2.2-I2V-A14B-Moe-Distill-Lightx2v by lightx2v, produced using convert_to_quant by silveroxides. Some layers have been kept on their original BF16 format, while others were quantized as MXFP8 or NVFP4, mostly.
Wan2.2-I2V-A14B-Moe-Distill-Lightx2v is an image-to-video generation model built on Wan2.2-I2V-A14B. It applies step distillation and a MoE architecture to reduce inference to 4 steps without CFG, cutting generation time substantially while preserving output quality.
Remember that you need to download the high and low noise models in order for your Wan 2.2 workflow to work.
IMPORTANT
Since NVFP4 is only supported on NVIDIA Blackwell architecture GPUs, running this model requires a Blackwell GPU with its corresponding support enabled in torch, along with a recent version of ComfyUI and comfy-kitchen built against CUDA 13.
Description
FAQ
Comments (10)
ELI5: what is this good for?
Good for people with low resources with an entry-level 50xx series GPU. Since this model uses NVIDIA's MXFP8 and NVFP4 formats, inference is faster than other quants, such as GGUF.
My setup: 1 x 8-GB RTX 5060, 32 GB RAM. A 5-second video can be generated in about 90 seconds whith these models. I'd dare say that it's possible to run these models with 16 GB RAM and 8-GB RTX 50xx GPU, provided you have a quick NVME disk with enough swap.
Also, the quants were not made blindly converting all possible layers to NVFP4 (4-bit quantization), such as other NVFP4 models hosted on Civitai. An analysis was run on the original models first to decide which layers where to be preserved as BF16 and which ones to be quantized using MXFP8 (8 bits) and NVFP4.
Do you have any comparison videos? I'd test myself but I'm away from my PC until next week. Curious how it would compare
Unfortunately I lack the resources to perform comparisons with the original model. Are you referring to comparisons with other NVFP4 quants being published here on Civitai?
I have an RTX 5060ti 16GB graphics card, 32GB of RAM, and the GGUF Q4KM compression mode takes up 15.5GB of video memory when running. Could you give me a rough estimate of how similar your model is to the GGUF Q5 in terms of quality and memory usage? Also, are there any limitations regarding compatibility with LoRa?
Which GGUF Q4_K_M compression mode are you referring to? These are NVFP4 quants for both high and low noise models of Wan 2.2 I2V from LightX2V. I'm using these models on my 8-GB RTX 5060 without memory issues. I'm using GGUF quants for the text encoders to save up some VRAM.
Regarding the compatibility with LoRAs, I cannot say, since I don't have enough resources to use them. My guess is that maybe, being these models quantizations of already distilled models some LoRAs won't be so effective, but you can always try.
I'm using a laptop with an RTX 5070 Ti, 12GB VRAM, and 48GB RAM. In practice, NVFP4 offers faster generation speed, with almost no noticeable difference in quality compared to Q4. Lora compatibility is not an issue—anything that works with 4-step Q4 also works with this.
NVFP4 is roughly 150% to 200% faster than Q4, and it consumes less VRAM and memory.
is this model will work in rtx 4060ti 16gb?
It might work provided you have comfy-kitchen support enabled and CUDA 13.x with Torch 2.10, at least. For the 40xx series, NVFP4 will get upscaled to FP8 or FP16 (not sure about that) and you won't get the benefits that the 50xx series cards have. I don't recommend you to use my models, to be honest.
You'll get much better results in terms of speed and quality with FP8 models or smaller quants from other users in Civitai, such as those from @Darksidewalker. Bear in mind that NVFP4 is a 4-bit format that's only supported on Blackwell devices and later.
To be precise, my models keep some layers on BF16, others on MXFP8 and others get converted to NVFP4. MXFP8 is another new format not supported on 40xx series, either.
Why do all videos made with Wan2.2 NVFP4 checkpoints look like 0.75x slow motion
Details
Available On (1 platform)
Same model published on other platforms. May have additional downloads or version variants.