This is a model that has been created and trained on all completely prompt generated AI images to portray the character whose name is Sebastian. He is a 19-year-old musician who dropped out of School . One of several original characters in a visual novel titled counterpoint.
Description
Still doing preliminary tests on this one.
Comments (4)
I took a look at your training data and it looks like the images were tagged rather than captioned in natural language.
The modern models like Z-Image, FLUX, Qwen-Image, Krea 2, and other newer image models are trained primarily on descriptive natural-language captions and use much stronger language understanding than older CLIP-centric workflows like SDXL and its variants (Pony, Illustrious, etc.). While tagged datasets will still work, you'll generally get significantly better prompt adherence and more natural, consistent results by training your LoRA with high-quality captions that match how the base model was trained. Just my two cents.
hey thank you so much for letting me know!!! yeah I’m gonna do that in my next one. So do I just bench the tags and describe each one? like what is happening? or is there an auto feature like autocaption or something.
@paintingwithbrooks377 @paintingwithbrooks377 Sure thing. I would definitely auto-caption unless you really like writing - and even then I wouldn't recommend it because it's better to have a similar approach on description syntax style and word choices that the base model was originally trained on. Qwen VLMs are pretty good at this and there are some easy tools to do it like https://civitai.com/models/2459492/qwen3-8b-vl-imagevideo-caption-uncensored
If you're using this platform to train then there's a caption tab next to the tag tab and it'll auto-caption that way.
Also, just like the old way with tagging, you need to check the work and make spot edits. Sometimes it'll mix up genders, camera angles, clothing types, etc, so make the necessary corrections before you feed it to the training engine. I've had a lot of good datasets make poor LoRAs by not double checking what it wrote down.
@paintingwithbrooks377 You also may want to do a little experement yourself and just retrain the dataset you used here but the only thing different is one LoRA trained with tags and the other with captions, and then give them a bake off test with various prompts from simple to really complex and you'll see the difference.


