This is HQ Int8 Row ConvRot of Krea 2 Raw (Base slow) model.
Made from official BF16 model with SECourses Musubi Trainer Quantization app
You can download and use Musubi Trainer app for both training and quantization from here : https://www.patreon.com/SECourses/posts/secourses-musubi-137551634
To be able to use this model with very best performance please use our Torch 2.13 CUDA 13 ComfyUI installer with ready presets : https://www.patreon.com/SECourses/posts/download-comfyui-installers-and-presets-105023709
I also recommend our SwarmUI installer with ready SwarmUI presets : https://www.patreon.com/posts/download-swarmui-installer-and-presets-114517862
Int8 ConvRot is 96.2% similar to BF16 meanwhile GGUF Q8 is only 90.0% and FP8 Scaled is 82.2% and NVFP4 is 63.7%
Moreover, Int8 ConvRot generates the output in 3.05 seconds, making it 1.82× faster than BF16, which takes 5.56 seconds.
NVFP4 takes 3.8 seconds and is 1.46× faster than BF16, whereas GGUF Q8 takes 6.06 seconds and is approximately 8.3% slower than BF16.
So Int8 ConvRot generated with our Musubi Trainer app at high quality is almost 100% faster and almost same quality as BF16
High quality generation takes few hours on RTX 5090
With our ComfyUI backend, Int8 Row ConvRot is able to generate faster than FP8 Scaled literally 100% faster on RTX 3000, 4000 and 5000 series GPUs
Model quantization is taking around 3-4 hours on RTX 5090 since we do training like quantization with prodigy optimizer
Check model screenshots to see and learn more
Description
This is HQ Int8 Row ConvRot of Krea 2 Raw (Base slow) model.
Made from official BF16 model with SECourses Musubi Trainer Quantization app
You can download and use Musubi Trainer app for both training and quantization from here : https://www.patreon.com/SECourses/posts/secourses-musubi-137551634
To be able to use this model with very best performance please use our Torch 2.13 CUDA 13 ComfyUI installer with ready presets : https://www.patreon.com/SECourses/posts/download-comfyui-installers-and-presets-105023709
I also recommend our SwarmUI installer with ready SwarmUI presets : https://www.patreon.com/posts/download-swarmui-installer-and-presets-114517862
Int8 ConvRot is 96.2% similar to BF16 meanwhile GGUF Q8 is only 90.0% and FP8 Scaled is 82.2% and NVFP4 is 63.7%
Moreover, Int8 ConvRot generates the output in 3.05 seconds, making it 1.82× faster than BF16, which takes 5.56 seconds.
NVFP4 takes 3.8 seconds and is 1.46× faster than BF16, whereas GGUF Q8 takes 6.06 seconds and is approximately 8.3% slower than BF16.
So Int8 ConvRot generated with our Musubi Trainer app at high quality is almost 100% faster and almost same quality as BF16
High quality generation takes few hours on RTX 5090
With our ComfyUI backend, Int8 Row ConvRot is able to generate faster than FP8 Scaled literally 100% faster on RTX 3000, 4000 and 5000 series GPUs
Model quantization is taking around 3-4 hours on RTX 5090 since we do training like quantization with prodigy optimizer
Check model screenshots to see and learn more
FAQ
Comments (16)
Is this basically an --int8 --scaling_mode row --convrot with convert_to_quant and leaving the simple flag off with default SVD quantization?
Face ID with Reactor Included ? Serious ? No Thank You.
@Weaze
What do you mean and where did you read that?
@vladulidlo ye what he says makes no sense lol
@vladulidlo In the description after klicking on the Patreon Link which is obvious changed now. Edit: No its not changed. klick on the second link in the description. be careful with this custom comfyui installer. reactors nfsw scanner and telemetric data collector is a massive security risk in my opinion which AI confirms.
Would you mind making one of raw? Since many people use Raw + Turbo lora combo.
@kossan I wonder why people cannot read anymore. It litreally writes Krea 2 Raw / Base.
@vladulidlo my bad, I was actually looking at the images first and figured that would be the actual version due to the comparisons being made with it.
@vladulidlo Int8 quantization is not raw.
@aising23 Raw is a non-turbo model aka base, not a quantization. Both is written in the title, so no matter what you mean, both is stated.
@vladulidlo The title also writes Int8 Row ConvRot, and the file size confirms it's not raw. SECourses also explains it's based on raw.
A raw model is stored in 16-bit Floating Point, and Int8 takes 16-bit decimals from the Raw model and rounds them down into simple 8-bit integers (whole numbers), cutting the file size in half. This is the mathematical quantization. Rounding decimals to whole numbers creates massive mathematical errors, resulting in quality loss or blank images. The ConvRot matrix prevents the loss.
Regarding Turbo, it can be injected to both raw and Int8 models.
@aising23 We aren't talking about that, so I'm not sure what the comment is about. We aren't talking about the file being raw, we are talking about the version of the model. There are two, Raw or Turbo.
He isn't claiming the model version he is using is the raw base, he is saying the model is the version designated as Raw.
@kossan But you were, and I believe you and others are confused. There could be many format versions based on a raw model. There are several formats a raw can be converted into, and Turbo is just another layer on top of these formats. Turbo can be added as a LoRa node or baked into formats.
The description suggest a raw model, then converted into Int8 ConvRot format. It is the "Type: checkpoint trained" on the right side that gets many confused. If it is a truly raw finetune (not LoRa stacking), such raw model is useful. If offered.





