v10
注意:这个版本尝试额外添加了一个触发词“ppw”。把它放在提示词的开始处来触发lora。
Note: This version experimentally adds an extra trigger word: “ppw”. Place it at the beginning of your prompt to trigger the LoRA.
降低了写实照片的训练轮数,尽管对于族裔标签的响应效果弱于v9,但是提升了生成质量。
In this version, I reduced the number of training epochs for the photo dataset. Although the response to ethnicity tags is weaker than in v9, the overall generation quality has improved.新增一个族裔标签Anime。使用它可以让人物获得更接近传统二次元的脸型。
Added a new ethnicity tag Anime. Using this tag can give characters a face shape closer to traditional anime.移除了质量提示词“low quality”的训练数据集。
The training datasets for the quality prompt tags "low quality" have been removed.对过去的数据集进行了维护。降低了学习率,增加训练步数,缓解了一些人物解剖错误、肢体扭曲的问题。
The version includes a maintenance of the previous dataset. The learning rate was reduced and the number of training steps was increased, which helps alleviate some anatomy errors and limb distortion issues.
快速上手 | Quick Start
这是什么? | What is this?
People's Works是一个实验性的微调模型系列,最初基于一个由Pony V6 XL生成的图片构成的数据集。这个模型的数据集由数千张AI社区用户发布的图片和作者使用AI合成的图片构成,经过人工筛选、编辑、修改和标注后用于训练。除此之外还有一个超过2000张由照片、游戏CG和3d渲染图像构成的辅助数据集,用于额外补充知识。
People’s Works is an experimental fine-tuned model series, originally based on a dataset composed of images generated by Pony V6 XL. The dataset consists of several thousand images published by AI community users, along with images synthesized by the author using AI. These images were manually curated, edited, modified, and annotated before being used for training. In addition, there is an auxiliary dataset of over 2,000 images composed of photographs, game CGs, and 3D renders, used to provide supplementary knowledge.
模型功能 | Model features
本系列模型主要功能是帮助基础模型在不使用画师串、较少的质量提示词的条件下获得相对稳定的风格化图像,为提示词节约token空间。
使用人工挑选和标注的数据集强化训练正面和负面审美提示词。
模型针对flat color、realistic等特定风格肌理进行了强化。
对人物的年龄、族裔、肤质等特征的精细化控制。
使用了较高的训练分辨率,使基础模型在高清修复时表现更好。
通过手工修复图片的方式降低了模型的某些细节瑕疵出现的概率。
The primary function of this series of model is to help the basemodel generate relatively stable, stylized images, without artist keywords or long quality tags, freeing up token space for prompts.
By using manually selected and annotated datasets, the model strengthens both positive and negative aesthetic tags training.
The model includes targeted enhancements for specific visual textures and styles, such as flat color and realistic.
It enables finer control over character attributes, including age, ethnicity, and skin texture.
A higher training resolution is used, improving the basemodel’s performance during high-res upscaling.
By manually editing the images, the likelihood of certain flaws appearing in the model’s outputs has been reduced.
使用方法 | Usage
v10
positive:
ppw, masterpiece, best quality, very aestheticnegative: /
v7-v9
positive:
masterpiece, best quality, very aestheticnegative:
low quality, displeasing更新记录 | Change log
v9
更改了系列名称。自这个版本起,训练集中来自Pony v6 XL的图片已经在所有AIGC内容中占比不足1/3。随着越来越多的新模型出现,我有计划在明年将这个系列拓展到其他模型上。为了避免未来用户使用上的混乱和误解,这个系列从这个版本起更改命名。
The series name has been changed. Starting from this version, images sourced from Pony v6 XL make up less than one third of the training data across all AIGC content. As increasing number of new models are emerging, I plan to expand this series to other models next year. To avoid potential confusion and misunderstanding for users in the future, the series name has been changed starting from this version.这个版本的训练方式是直接训练LoCon,而非训练Checkpoint后再抽取LoRA。模型相较前一个版本效果更强。
This version is trained directly as a LoCon, rather than training a checkpoint first and then extracting a LoRA. Compared to the previous versions, the model delivers stronger effects.v9全系列使用1536分辨率的训练集。现在使用这个Lora生成图片时支持单边768-1536的分辨率,使用高清修复时也可以尝试更高的denoise参数了。
All v9 models use a 1536-resolution training set. When generating images with this LoRA, single-side resolutions from 768 to 1536 are now supported. When using high-res fix, you can also try higher denoise values.对训练图片调色。现在模型在没有指定色彩时更倾向于生成暖色调的图片,并且色彩的饱和度略微提高。较暗的场景明暗对比更大了。
Color adjustments were applied to images. When no specific color is specified, the model now tends to produce warmer tones, with slightly increased saturation. Darker scenes also have stronger contrast between light and shadow.之前版本的数据集中,人物鼻子的画法不统一。出于作者本人的兴趣,手工修改了其中约300幅图片,并暂时排除了约200幅来不及修改的图片。现在人物的鼻子有鼻翼了。
In earlier versions of dataset, nose depiction was inconsistent. Out of personal interest, the author manually modified around 300 images and temporarily excluded about 200 images that could not be edited in time. Characters’ noses now have nose wings.删除了旧版本中数百张过时的低质量训练数据。
Hundreds of outdated, low-quality images from older versions of dataset have been removed.添加了一个新的实验性数据集: 使用真人相片作为引导,现在你可以使用以下年龄和族裔标签了:
A new experimental dataset has been added. Using real photographs as guidance, you can now use the following age and ethnicity tags:
child, teenage, adult, matureCaucasian, Asian, Indian, African我的标签设计优先选择Danbooru数据集中已经存在的标签,尽管其中很多只有很少量的数据,在原版模型中几乎无法触发。重新启用了已经被删除的danbooru词条Caucasian和teenage,增设了adult和African两个标签。此外,loli和shota因为其文化背景中强烈的性暗示倾向,这两个词条被完全替换,根据具体情况分流入child和teenage。
My tag design prioritizes labels that already exist in the Danbooru dataset, even though many of them had very limited data and are therefore almost impossible to trigger in the base models. The previously removed Danbooru tags Caucasian and teenage have been re-enabled, and two new tags, adult and African, have been added. Additionally, due to the strong sexual connotations of loli and shota in their cultural context, these tags have been completely replaced and redistributed into child and teenage depending on the situation.
v8
肌理更新:强化了以下tag的学习:
Texture Update: The following tags have been reinforced in training:
realistic, photorealistic, flat color,shiny skin, matte skin, shiny hair,请注意,在danbooru数据集中有很多个用于描述“照片”和“接近照片的风格”的tag。我在训练集中统一标注这些图片为“photorealistic”。但是使用danbooru训练集训练的SDXL模型大多并不能很好地画出写实图像,因此“photorealistic”只建议在较小权重下用于改变画面的肌理。“realistic”可以在高权重下正常运作。
Please note that Danbooru dataset contains multiple tags to describe "photo" or "photo-like styles". I’ve tagged all such images as “photorealistic” in dataset.
However, most SDXL models trained on the Danbooru dataset do not render realistic images well. “photorealistic” is only recommended at low weight, where it can help adjust texture rather than create realism images. The “realistic” tag can work properly at higher weight.
v8版本简介:
Pony: People's Works (ppw)是一个实验性的微调模型系列,数据集有约85%是收集自CivitAI上用户发表的AI生成图片。早期ppw的数据集最初建立在由pony v6生成图片的基础上,因此本系列模型生成的图片也带有pony diffusion的特征。
本系列模型使用标准Danbooru标签,主要擅长生成中、近景风格化人像。它们的主要功能是使基础模型可以在不使用画师串、较少的质量提示词的条件下获得相对稳定的图像质量,为提示词节约token空间。
本模型并非风格模型,在不同的提示词和生成条件下可能会有微妙的画风差异。
Pony: People's Works (ppw) is a experimental fine-tuned model series, approximately 85% of the dataset comes from AI-generated images published by users on CivitAI. Since the earlier ppw dataset was built on images generated by Pony V6, the outputs of this series also carry some characteristics of Pony Diffusion.
This series uses standard Danbooru tags and is mainly optimized for generating stylized portraits at medium and close range. The primary effect of this model series is to allow the basemodel to achieve relatively stable image quality, without artist keywords or long quality tags, freeing up token space for prompts.
These models are not style LoRAs. There may be subtle stylistic variations depending on different prompts and generating conditions.
v7
v7版本对数据集结构做了较大幅度的调整,并且使用了不同的训练参数和训练策略,因此可能v7会不如原来的版本稳定。
The v7 version has undergone significant structural adjustments to the dataset, and utilizes different training parameters and strategies. As a result, v7 may be less stable than the previous versions.
v-pred模型在civitAI在线生成器上的表现和吐司的在线生成表现完全不一样,相同参数完全无法复现。我也不知道是为什么......
The v-pred model's performance on the CivitAI online generator is completely different from online generation on TensorArt. The results are entirely unreproducible with a same parameters. I have no idea why...
TensorArt version CivitAI ver. with same parameter on CivitAI with higher weight
v7版本简介:
这是一个在前作的数据集基础上发展而来的图像质量LoCon,约90%-95%图片数据来自CivitAI上发布的图片。
它使模型可以在不使用画师串、使用较少的质量提示词的条件下获得相对稳定的图像质量,节约出更多的token空间,同时它还可以修复一部分模型固有的生成瑕疵。(但是不包括手部)
因为数据集选取的原因,生成的图片会带有Pony的质感。但因为它并不指向任何特定的画师、风格和绘画技法,所以在不同的提示词、模型条件下可能会有微妙的画风差异。
This is a generation quality LoCon developed based on the dataset from the previous work. About 90%-95% of the image data comes from CivitAI.
It allows models to achieve relatively stable image quality without artist tags or using long quality prompts, freeing up more token space. Additionally, it can fix some inherent generation flaws of the model. (except for hands)
Due to the dataset selection, the generated images exhibit a Pony-like style. However, since it does not reference any specific artist, style, or painting technique, there may be subtle stylistic variations depending on different prompts and checkpoint conditions.
数据集来源及许可证 | Dataset Source & License
数据集中每一张图片都经过作者本人的人工筛选、分类和标注编辑,其中上千张图片经过人工的编辑、对细节瑕疵进行修正。
此模型为免费、开源模型,用户可以在私人设备上自行部署该模型。作者并不从模型出售中获取任何报酬。作者并不限制本系列模型用于商业生成服务或者生成图像用于商业用途,但是请注意配合使用的Checkpoint和其他LoRA的许可证限制。
请注意本模型的数据集由一个为约5000张AI图像构成的训练数据集和一个超过2000张图像构成的辅助数据集组成。其中,主数据集的图片大部分收集自AI社区。辅助数据集则包括公开的新闻图片、游戏CG和宣传图、3D渲染图和已购买的商业写真等。辅助数据集仅用于学习光影色彩、构图术语、人体解剖特征等一般通用知识,使用本模型不能还原辅助数据集中的版权内容。当前法律对这类数据的使用没有明确的统一规定,请有商用意向的本系列模型用户自行注意相关风险。
本数据集没有训练任何独立画师的数据,也没有标注任何画师ID信息(不排除AI错误标注的情况)。
另外,本模型不允许用作闭源商用、模型出售,也禁止用于闭源商用模型的融合。对于开源融合模型用于生成服务的情形不做限制,但是建议标注融合模型的出处。
Every image in the dataset was manually reviewed, categorized, and annotated by the author. Among them, more than a thousand images were additionally hand-edited to manually correct fine-detail visual defects.
This model is free and open-source model, allowing users to deploy it on their personal devices. The author does not receive any compensation from selling the model. The author does not impose restrictions on using this model for commercial image generation services or generating images for commercial purposes. However, please be mindful of the license restrictions of the Checkpoint and other LoRAs used alongside this model.
Please note that this model’s dataset consists of a primary training set of approximately 5,000 AI-generated images and an auxiliary dataset of over 2,000 images. The majority of images in the main dataset were collected from AI communities. The auxiliary dataset includes publicly available news photographs, game CGs and promotional images, 3D-rendered images, and purchased photo sets.
The auxiliary dataset is used solely to learn general-purpose knowledge such as lighting, composition terminology, and human anatomical features. This model cannot reproduce or restore copyrighted content from the auxiliary dataset. As current laws do not provide a clear, unified standard for the use of such data, users who intend to use this model for commercial purposes should be aware of and assess the associated legal risks on their own.
This dataset does not include training data from any individual artist, nor does it contain explicit artist attributions (though AI mistagging cannot be entirely ruled out).
Additionally, this model is not permitted for use in closed-source commercial applications, model resales, or merged into closed-source commercial models. There are no restrictions on open-source merged models being used for image generation services, but it is recommended to credit the sources of any merged models.
Description
FAQ
Comments (19)
关于v9的版本:
about the LoRA variants in v9:
这个版本的模型是本系列开始以来,数据集变动幅度最大的一个版本。因为测试的次数过多,训练成本超过了2500元(近400美元),对于我个人来说有些超支。所以我暂时只打算发布这四个使用量最大的版本的模型。
Illustrious系列目前流量最大的模型毫无疑问是v1.0。除此之外,我发现以Hassaku v1.3 style A为代表的基于Illusv0.1的模型使用量也不低。相反,几乎很少看见有人以Illus v2.0作为底模进行训练。尽管people's works v8 基于Illus v2 Stable训练的lora使用量最高,但是实际上大都是应用在基于Illustrious v1.0训练的Basemodel上。所以我这个版本把所有版本的训练分辨率都上调至了1536,相应地,又去掉了Illus v2的训练,添加回了Illus v0.1的版本。这样应该可以最大化生成质量。
另外,我注意到NoobAI有一个新分支CKXL,但是我训练ppw v9时,CKXL只发布了v0.1。我会持续关注这个系列,并且下个版本可能会添加支持。
上个版本在一些用户的建议下添加了对Rouwei的支持,但是实际上发布后几乎没有人使用rouwei生成。所以这个版本我打算暂停支持。我关注到作者正在引入新的vae和te进行训练,我会观望这个模型的下一步发展(但是很不幸,它正巧遇到了z-image这个对手,所以这个研究有些前途未卜)。
Animagine v4 本身并不是一个差劲的模型,但是可惜生不逢时,几乎没有人使用它。Pony v6在性能上已经完全不敌今年的新模型,我以后也很可能都不会再为它训练Lora了。
This version has the largest dataset changes since the beginning of this series. Due to excessive testing, the training cost exceeded 2,500 CNY (nearly USD 400), which is over budget for me personally. Therefore, for now I only plan to release these four model variants with the highest usage.
Within the Illustrious series, the model with the highest traffic is undoubtedly v1.0. In addition, models based on Illus v0.1 (represented by Hassaku v1.3 style A) also see considerable usage. In contrast, it is rare to see people training with Illus v2.0 as the base model.
Although v8 trained on Illus v2 Stable has the highest LoRA usage, in practice, most of the generations are applied to basemodels trained on Illustrious v1.0. Therefore, in this release I have increased the training resolution of all variants to 1536, removed Illus v2 from training, and added back Illus v0.1. This should maximize generation quality.
I also noticed that NoobAI has a new branch called CKXL, but when I trained PPW v9, CKXL had only released v0.1. I will continue to monitor this series, and future versions may add support for it.
In the previous release, support for Rouwei was added based on user suggestions. However, after release, almost no one actually used Rouwei for generation. As a result, I plan to pause support in this version. I’ve seen that the author is introducing new VAEs and TEs for training, so I will observe the model’s future development (though unfortunately, it happens to be facing a competitor z-image, which makes the future of this research uncertain.).
Animagine v4 itself is not a bad model, but it was released at an unfortunate time and sees almost no user. Pony v6 is already completely outclassed by this year’s newer models, and I will most likely not train LoRAs for it in the future anymore.
接下来一段时间的工作安排:
Work Plan for the Coming Period:
Because the training resolution of this version has been increased, I removed a large number of blurry images and images with noticeable detail flaws (especially by pony v6). As a result, the basic style of v9 is somewhat unstable. The main focus in the near future will be synthesizing high-resolution data to ensure style consistency.
The new ethnicity feature is an experimental attempt and still requires much more data. Both the Danbooru dataset and the original SDXL model show strong bias toward the “Asian” tag, resulting in images with a pronounced aesthetical orientalism, which significantly contaminates training results. I will spend some time trying to improve this. In addition, data for many ethnic groups is difficult to collect. This feature is still in an early development stage, and the current dataset only supports female subjects, so there is a lot of work to be done going forward.
With the rapid rise of DiT models this year, it is likely that many anime-focused finetuned models based on DiT will emerge next year. I will start preparing natural-language datasets for training in the future.
Training v9 took me five months, making it the longest release cycle in this series so far. My ideal release cadence is three months, or once a season. If all goes as planned, the next version should be around March.
因为这个版本的模型提升了训练分辨率,我删除了大量模糊不清和细节有瑕疵的图片。这导致了v9版本的基础画风有一些不太稳定。接下来一段时间的主要工作是合成高分辨率的数据,保证图像的风格稳定。
新版本的ethnicity功能是一次新的尝试,我还需要更多的数据。Danbooru数据集和SDXL的原始模型都对Asain这个tag有很强的bias,生成的图片带有浓厚的东方主义色彩,这对训练效果造成了很大的污染。我会花一些时间尝试改进。此外,很多族裔的数据都有些难收集。这个功能还处于开发早期,而且目前数据集只支持女性,未来还要做很多工作。
随着今年DiT模型的大爆发,明年应该会有很多基于DiT的二次元微调出现。我要开始着手准备自然语言数据集了。
v9的训练花费了5个月,是这个系列最长的一次周期。我理想的发布周期是3个月,也就是一个季度一次更新。我想下一次发布应该是在3月份。
Love this
Not sure why but when i use v9 i get purple noise explosions and have to reload the model
Hello. I haven’t encountered this issue when I tested them on ComfyUI.
Would you mind sharing a bit more information? Which UI are you using? If you load only the checkpoint without any LoRA, does it run normally with the same settings? Under the same configuration, does this issue occur when using other LoRA or LyCORIS? Does this issue happen on only one v9 model, or does it happen on all v9 versions? And does this happen when using v8 or lower version?
@Dajiejiekong I can actually tag in; in my case, im on SD.next, so a branch of automatic1111. All versions of v9 does it regardless of E or V-pred, V8 and below are fine. Im suspecting it have something to do with the webui itself vs Locon/Lycoris. There are few other Loras that have the same problem in my testing, namely Genesis 2.0 and Phad v1.0. Idk if there's anything you can do to help tho.
tldr: 麻麻我不想学comfyUI QAQ
@iboter 问一下哈,你的WebUI里有没有一个叫做LyCORIS的插件?
@Dajiejiekong 没有,因为插件会和Webui有冲突。 我花了了一个小时学Comfy现在把SDnext删了.. 笑
测了一遍上面举例的,应该就是WebUi本身问题。装插件会有冲突,LyCORIS放自己文件夹或LORA都不行。Comfy全都可以跑。 按理来说24年的更新应该把Lora和Lycon整合了所以插件废用了但不知道为什么检测不到。用Original backend可以但1024x1024的图能跑半小时..
Webui 燃尽了
@iboter 我开始的猜测就是插件冲突,既然没有那就不是这个问题了。按理说应该放LoRA文件夹就可以的。因为我平时用的笔记本,跑不了SDXL,画图都是租卡或者用在线服务,所以测试图全是ComufyUI出的,没有想到过会出这个问题。
这个系列从v4版本开始就是LoCon模型,按理说本身这个东西已经出现好几年了,不应该出这些问题才对。唯一的区别就是v1-v7是用秋叶训练器,v8和v9是用的kohyaGUI训练的。v8是训练大模型后抽取的,v9是直接训练的LyCORIS。
但其实这两个都是UI,实际上运行的都是kohya_ss这个训练脚本,按理说结果也不应该有什么区别才对。那我觉得要么是kohyaGUI这个UI有什么bug,要么就是kohya_ss最近半年的更新变动了什么。也可能是某些WebUI的分支本身和这些脚本中的某些东西有冲突。这个就很难说了。
不过因为这些UI平台都有些年久失修,后面新的模型架构也大概率不支持,我可能接下来也会换非kohya_ss的训练器了。
really good
perhaps one of the best LORAs on the site.
awesome!
Please explain to me who came up with the idea to set -0.7 for this lora? And why?O_O
what strength is recommended to run this at?
0.9-1

