Update v2:
retrained with 2 changed variables: lr 0.0002 and much more detailed captioning with qwen3-vl. its much better alsthough there is some overfit, but it's fine maybe?!?
the right settings/ dataset for z-image are still a mystery to me.
i now suspect it really wants detailed captions to work well.
don't know if the overfit is solvable.
but anyways: works much better now and seems really compatible with style or character lora....
v2 and v2.12 are just different epochs. as i can't for the life of me decide which ones better.
V1 notes: Still not too sure about z-image training...
results seem to be 50/50 good bad.
training notes v1: small dataset with portrait selection (90 Images), 1600steps, 0.0005, 32/32
Description
FAQ
Comments (3)
the problem with zit loras is the distilled model. HF has a de-distilled checkpoint someone made to mitigate it until they release the base model
i know about that and that it probably causes the overfit, i still want to figure out what settings to use for this model (there is a lot of improvement between v1 and v2 of this lora with the same dataset),because it's the model we have. it still is in most cases better than anything else i've used (with comparable speed)
v2.12 looks good for me. It works very well with 0.4 str. Thank you for it.
