nvfp4 quant created with https://github.com/silveroxides/convert_to_quant tool
Description
FAQ
Comments (20)
Finally! been waiting for an fp4 version and Just to clarify, this is just the transformer?
not just one transformer. a video\audio vae decoder\encoder and vocoder layers included. this layers and a part of transformer layers still in bf16 precision.
@DooFY87 Yeah figured it was after and have been using it, so it should be 'mixed' precision? Strange how it's smaller than my nvfp4 mixed version by Hippotes
Thanks anyway sir!
@MrTitsworth Maybe Hippotes used different tools or conversion settings, so that could explain the size difference.
Does this require a 5000 card?
No, you can use it on any card you just won't get the specific speed benefit that nvfp4 provides however if you have low vram/ram it will be faster than heavier models so it's a bit of a win still for low resource systems that are not 50xx (I use it on my 3070ti and it's just as fast as INT8 however quality is a bit less but more stable than int8 for me)
If you've got lots of VRAM/RAM tbh I'd skip FP4 but if not it can be good.
i tested this on 40xx card, all working ok, but without 50xx cards specific speedup
How much faster is it on a 5090, and is it worth the quality drop?
speed is the same. Just faster loading
NVFP4? Or FP8/MXFP?
I look inside. I see FP8 and BF16. =)
I want to try to quantize the LoRA into a similar quant. I want to know exactly what to choose in silveroxides/convert_to_quant.
Check with Google regarding the technical specifications of NVFP4 quantization and how it's packed. This will explain why we can not see a "NVFP4" inside converted model.
I haven't attempted LoRA conversion myself. However, convert_to_quant includes a help interface via the -h argument with additional sub-sections; you might find the information you need there.
судя по нику, можно и на другом языке ответить: все nvfp4 значения пакуются в fp8, по 2 штуки, это особенности формата от nvidia) их официальные конвертации различных моделей выглядят также. лоры я не конвертировал, можно глянуть справку, запустив convert_to_quant -h, может что найдётся
@DooFY87 А в этой штуке, которая convert_to_quant выбирать надо --nvfp4 флаг, я верно понимаю? =) По сути, мне это главное, просто для душевного равновесия, чтобы были одинаково сконверчены. Ну вот и попробую. ^_^' Что получится.
@BahamutRU верно, --nvfp4)
What ComfyUI node do you use to load the dit (transformer) from this safetensors file? Is it LTX2_SM_Model with the dit parameter, or UnetLoaderGGUF, or something else? And is there a specific version/branch of the node needed?
This not simple transformer model, VAE and some other stuff included. I using standard nodes: Checkpoint loader, LTXV Audio Text Encoder Loader, LTXV Audio VAE Loader. You only need additional gemma 3 text encoder model
Could you please share the workflow you use? I'm new to this.
Use LTX 2.3 workflow from standard comfyui templates as example, remove some unneeded models if this present inside workflow. I using standard nodes: Checkpoint loader, LTXV Audio Text Encoder Loader, LTXV Audio VAE Loader. You only need additional gemma 3 text encoder model
This model does not wash out by twenty seconds of generation, or does it so much less than all other LTX checkpoints that I can't notice.
I really want to stress this point. Other models's gens, quantized or not, start "dissolving" in quality past fifteen seconds, maybe twelve.
Or I was lucky during the past thirty gens. Either way, impressive model. Thanks so much ! THIS is the one that made me leave WAN behind!
Now, if we just could have better voice actor control...
Could someone tell me if 14.7 gb of Vram is enough to make videos of intermediate or low quality?