My primary intent of this quant is to make it run within 32 GB RAM while still having enough memory to run some of my other tasks without performance degradation. If quality is your priority and hardware is not an issue, I would recommend seeking INT8 Convrot (~14 GB) or better
Experimental quantization of Kroma v0.2 by Lodestones, primarily aimed at reducing VRAM usage and improving usability on older consumer GPUs.
This build uses a mixed quantization approach with INT8 tensor-wise quantization and W4A4 layers.
v0.1 is kind of that sweet spot I found, not too destructive in terms of quality
v0.2 and v0.3 are further smaller is size and of course relatively lower quality (will upload em if anyone actually wants them)
Used a simple T2I workflow, no second stage or upscaling done (I'm too lazy)
Tested on:
NVIDIA GTX 1660 Super 6 GB
32 GB System RAM
Test settings:
8 steps
CFG 1
Euler / Simple
696 × 1048 resolution
Approximately 8 s/it on a GTX 1660 Super
This is an experimental release, so performance and compatibility may vary depending on hardware and workflow.

