Not much to say about v2.0. It's superior in image quality, has better lighting and shading and tends more toward hyperrealism than realism itself. I'm still doing tests to see how far this turbo model can go, but I'm looking forward to the main model coming out. Very negative point about the v2.0: it's purposely overcooked and doesn't get along well with LoRas, unless you train them specifically for this version.
My conclusion two months after the release of this model is that the base model with LoRas is the best option, until they reveal (if they reveal it at some point, which I doubt) the weights to be able to make a full finetune.
It's only been a few days since the release of Z image Turbo, and I've liked the model so much that I haven't stopped testing all sorts of things to explore its versatility. I want to share my first results in the form of this first modified checkpoint, which aims to improve something that's already very good. I hope you'll try it out and share your images and impressions.
Description
v1.0
FAQ
Comments (13)
fp8 or gguf when?
FP8 is currently being uploaded. I haven't had time to quantize GGUF yet, but I will in a few days.
I want to know how does Z-Image compare to the various illustrious models and pony models out there? Do you think it's worth investing time in to use? And do you think there will be a robust LoRA community around it?
I'm not the creator but I'll try to give my own answer.
1) IMO, it's incorrect to compare ZIT with Pony/Illustrious. Pony/Illustrious are tuned towards anime, ZIT is a generalist checkpoint. Thus, ZIT should be compared with SDXL, not pony/illustrious. Pony/Illustrious are comparable with vanilla Chroma.
2) ZIT is a next-gen checkpoint, it uses LLM as a text encorder like FLUX2 and Qwen. But both FLUX2 and QWEN are monsters in terms of VRAM and disk space consumption. Also, personally I win nothing on my typical resolutions (rarely higher than 1152x1152) with Qwen/FLUX2 . ZIT provides the most balanced choice among next-gen checkpoint.
3) Personally, I believe, yes. The reasons I decribed in 2) give better opportunities to train loras with casual PC.
1) At the moment I think it is too early to say if it is worth investing time in fine-tuning this model. Mainly because at the moment they have not released information about full fine tuning and that the two other z image models that are estimated to still perform better have yet to be released.
2) As for whether the community will build a solid LoRa base around Z image turbo, I have no doubt that a lot of LoRa will be trained with this model, you can already find many after just a few days. When the base Z image model is released, we will see if the turbo variant remains just as popular or is it just the hype of the premiere. I imagine it will depend a lot on the hardware requirements of the base model Z image (or whatever they are going to call it).
3) Think about a future Z image turbo model, finely tuned in the style of Pony and illustrious? Too soon to know. I think it will depend on the fine tuning facilities of the model and especially those of the base model to be released. You can read many posts about how z image turbo is the spiritual successor of SDXL, but today in my humble opinion, I would wait to see the following developments. Getting to the point of tuning a model comparable to Pony or Illustrious is a big deal and requires a lot of time, dedication and money.
@DeViLDoNia @mphobbit I think I agree that it's too early to tell. And while Pony/Illustrious are tuned toward anime originally, the reality is with the plethora of custom merges and fine tuned models geared toward realism as well as related LoRAs, illustrious and pony models are my go to when trying to do realism. Perhaps Z-Image models will overtake illustrious/pony model usage over time.
Who knows. Couple of weeks ago I read everywhere how ZIT is bigger and better than FLUX. Couple of days ago people starting praising FLUX Klein though. Another time people swore Chroma will change it all. Let me give my honest opinion:
It could change by the day and there is no one who can provide you with a sufficient answer to your question.
If you want to try it without investing to much time, just take the ZIT Workflow from @SubtleShader (thanks Subtle!!). Either use their ZIT2 model or the one from the pornmaster dude on civit.
I like ZIT. Quick, great artistic quality and for that speed probably the best photo realism.
The sample images are looking great!
I'm very curious how did you actually train this checkpoint. Was it a bunch of LoRAs merged into the main checkpoint? Or was it a fine-tune? Did you employ a GUI trainer like AI Toolkit? How did you caption your dataset?
I've been slowly building a dataset consisting of high-resolution images, that was initially meant to enhance Flux 1's realism and skew it towards my preferred aesthetic (getting rid of that chin among other things 😄) but then came Qwen-Image so my focus changed, and now ZiT which is, from what I could see, an absolute banger of a realistic model and exactly what I was looking for as a base model.
My dataset now has thousands of images, and I wasn't sure what would be the optimal approach in training for realism.
It is a merge of several LoRa's. There is no fine tuning available yet, I have simply been testing everything a little. For captions I use a system with LLM, which is giving me good results. The number of images per lora varies a lot, I have not yet found the best number. But around 20 images with a good caption seems to be a win win. I have also achieved good and bad results with fewer images and with more. I actually use AI Toolkit at the moment, but it gives me some problems and I have already reinstalled it a couple of times.
You did a huge work bro. Amazing
aw*some
What sampler / steps / cfg settings are recommended?
Details
Files
Available On (1 platform)
Same model published on other platforms. May have additional downloads or version variants.

















