Anima is a 2 billion parameter text-to-image model created via a collaboration between CircleStone Labs and Comfy Org. It is focused mainly on anime concepts, characters, and styles, but is also capable of generating a wide variety of other non-photorealistic content. The model is designed for making illustrations and artistic images, and will not work well at realism.
It is trained on several million anime images and about 800k non-anime artistic images. No synthetic data was used for training. The knowledge cut-off date for the anime training data is September 2025.
Versions
Anima-Base
The pretrained, unrefined base model. Maximum flexibility, diversity, and style adherence.
LoRAs should be trained using this version.
Anima-Aesthetic
Fine-tuned for better consistency and a higher quality default art style.
Anima-Turbo
Distilled version for fast generations.
Use at CFG 1 and 8-12 steps.
The distillation process also increases stability and gives the model a strong default style, but reduces diversity.
I recommend starting with Anima-Turbo. On average, it is only slightly worse than Anima-Aesthetic, while being very fast to generate (and much cheaper if you use it on an online platform that scales the cost with step count). This makes it very convenient for quickly iterating on prompts. The increased stability can even make it better than Aesthetic in some cases.
Installing and running
Get the text encoder and VAE from the HuggingFage page.
The model is natively supported in ComfyUI. The model files go in their respective folders inside your model directory:
anima-base-v1.0.safetensors goes in ComfyUI/models/diffusion_models
qwen_3_06b_base.safetensors goes in ComfyUI/models/text_encoders
qwen_image_vae.safetensors goes in ComfyUI/models/vae (this is the Qwen-Image VAE, you might already have it)
Generation settings
Works at resolutions between 512^2 and 1536^2 pixels.
30-50 steps, CFG 4-6.
The Aesthetic version can tolerate lower CFGs such as 3, and often looks better with them.
A variety of samplers work. Some of my favorites:
er_sde: neutral style, flat colors, sharp lines. I use this as a reasonable default.
euler_a: Softer, thinner lines. Can sometimes tend towards a 2.5D look. CFG can be pushed a bit higher than other samplers without burning the image.
dpmpp_2m_sde_gpu: similar in style to er_sde but can produce more variety and be more "creative". Depending on the prompt it can get too wild sometimes.
euler: a basic sampler that is a bit more creative than er_sde. Good with the Turbo and Aesthetic versions, since those are naturally more stable.
If going for a more realistic / painterly look, the beta57 scheduler (ComfyUI RES4LYF custom node pack) can help make better textures, since it puts more emphasis on low-noise timesteps.
Prompting
The model is trained on Danbooru-style tags, natural language captions, and combinations of tags and captions.
Use lowercase for tags, and spaces instead of underscores. Score tags are the only tags that use underscores.
Recommended positive prefix: "masterpiece, best quality, score_7, safe, "
Recommended negative: "worst quality, low quality, score_1, score_2, score_3, artist name, blurry, jpeg artifacts, chromatic aberration"
When using a tag that is different between Danbooru and Gelbooru, prefer the Gelbooru version.
Prompt weighting works, but needs a weight higher than typically used for SDXL. Example: "(chibi:2)"
Aesthetic Version Prompting
Anima-Aesthetic is fine-tuned only on high quality images, with all of the quality tags stripped out from the captions. You don't need to use quality tags in the positive at all, but "masterpiece, best quality, " is safe to leave in. I recommend not using score_* tags in both the positive and negative prompt. It is already high quality enough and the score tags can push it too hard into slop territory.
Tag order
[quality/meta/year/safety tags] [1girl/1boy/1other etc] [character] [series] [artist] [general tags]
Within each tag section, the tags can be in arbitrary order.
Quality tags
Human score based: masterpiece, best quality, good quality, normal quality, low quality, worst quality
PonyV7 aesthetic model based: score_9, score_8, ..., score_1
You can use either the human score quality tags, the aesthetic model tags, both together, or neither. All combinations work.
Time period tags
Specific year: year 2025, year 2024, ...
Period: newest, recent, mid, early, old
Meta tags
highres, absurdres, anime screenshot, jpeg artifacts, official art, etc
Safety tags
safe, sensitive, nsfw, explicit
Artist tags
Prefix artist with @. E.g. "@big chungus". You must put @ in front of the artist. The effect will be very weak if you don't.
Full tag example
year 2025, newest, normal quality, score_5, highres, safe, 1girl, oomuro sakurako, yuru yuri, @nnn yryr, smile, brown hair, hat, solo, fur-trimmed gloves, open mouth, long hair, gift box, fang, skirt, red gloves, blunt bangs, gloves, one eye closed, shirt, brown eyes, santa costume, red hat, skin fang, twitter username, white background, holding bag, fur trim, simple background, brown skirt, bag, gift bag, looking at viewer, santa hat, ;d, red shirt, box, gift, fur-trimmed headwear, holding, red capelet, holding box, capelet
Tag dropout
The model was trained with random tag dropout. You don't need to include every single relevant tag for the image.
Dataset tags
To improve style and content diversity, the model was additionally trained on two non-anime datasets: LAION-POP (specifically the ye-pop version) and DeviantArt. Both were filtered to exclude photos. Because these datasets are qualitatively different from anime datasets, captions from them have been labeled with a "dataset tag". This occurs at the very beginning of a prompt followed by a newline. Optionally, the second line can contain either the image alt-text (ye-pop) or the title of the work (DeviantArt). Examples:
ye-pop
For Sale: Others by Arun Prem
Abstract, oil painting of three faceless, blue-skinned figures. Left: white, draped figure; center: yellow-shirted, dark-haired figure; right: red-veiled, dark-haired figure carrying another. Bold, textured colors, minimalist style.deviantart
Flame
Digital painting of a fiery dragon with glowing yellow eyes, black horns, and a long, sinuous tail, perched on a glowing, molten rock formation. The background is a gradient of dark purple to orange.Natural language prompting tips
Follow standard English capitalization rules for character and series names.
If using pure natural language, more descriptive is better. Aim for at least 2 sentences. Extremely short prompts can give unexpected results.
You can mix tags and natural language in arbitrary order.
You can put quality / artist tags at the beginning of a natural language prompt.
"masterpiece, best quality, @big chungus. An anime girl with medium-length blonde hair is..."
Name a character, then describe their basic appearance.
"Digital artwork of Fern from Sousou no Frieren, with long purple hair and purple eyes, wearing a black coat over a white dress with puffy sleeves..."
This is extra important when prompting for multiple characters. If you just list off character names with no description of appearance, the model can get confused.
Limitations
The model doesn't do realism well. This is intended. It is an anime / illustration / art focused model.
The model may generate undesired content, especially if the prompt is short or lacking details.
Avoid this by using the appropriate safety tags in the positive and negative prompts, and by writing sufficiently detailed prompts.
The model isn't great at text rendering. It can generally do single words and sometimes short phrases, but lengthy text rendering won't work well.
The base version is a true base model. It hasn't been aesthetic tuned on a curated dataset. The default style is very plain and neutral, which is especially apparent if you don't use artist or quality tags.
Finetuning tips
Don't train the LLM adapter. My own training script, diffusion-pipe, lets you set llm_adapter_lr=0 to completely disable training it, and the example config has this as a default.
Other trainers like sd-scripts have similar options that should be used.
The LLM adapter processes the text embeddings before they get to the diffusion model, and therefore has an outsized influence on the generated images. The adapter itself contains a surprising amount of knowledge and is easy to degrade by training it.
Use a low learning rate. For a rank 32 LoRA, start with 2e-5 and adjust up or down from there.
As a base model, there is no aggressive aesthetic tuning or RLHF you need to overcome when finetuning.
The model has an extremely large and diverse amount of visual concepts baked in already. A light touch is all you need.
Example of a style LoRA, with dataset and configs shared.
Online platforms
In addition to CivitAI, the following platforms also officially support Anima for hosted image generation.
License
This model is licensed under the CircleStone Labs Non-Commercial License. The model and derivatives are only usable for non-commercial purposes. Additionally, this model constitutes a "Derivative Model" of Cosmos-Predict2-2B-Text2Image, and therefore is subject to the NVIDIA Open Model License Agreement insofar as it applies to Derivative Models.
If you would like a commercial license, please email [email protected]
Built on NVIDIA Cosmos.
Description
Highres training is in progress. Trained for much longer at 1024 resolution than preview2.
Expanded dataset to help learn less common artists (roughly 50-100 post count).
FAQ
Comments (489)
Showing latest 282 of 489.
Any program for anima 3 training ? In lightning.ai
for lora look at the standalone anima trainer. for checkpoint it seems comfy? idk about training checkpoints
I love it! However it still doesn't understand certain concepts like: Licking, anal play (it knows some, but it could use a bit more), and some light domination like foot on back, foot in mouth, etc.
Keep up the great work!
FR. only and only downside of this model is it cannot understand insertion stuff.
Is it normal for the image to have artifacts after upscaling (1.5x) in img2img?
Forge Neo*
yes you can't Always fix stuff with these
Yep, same as PonyXL and Illustrious, waiting for training on larger resolutions
Both online and local generation deliver incredible results. Thank you so much for creating this fantastic model. If only you could agree on a license with Civitai to train LORAS models, Anima would be a resounding success. Best of luck, and I hope future versions are just as good. 💖
can i dm u , im having prob to gen img with it
Can I use it on A111!?
You can use it with forge NEO
I tried train multiple loras recently and found out there are definitely better than illustrious one! Great Work!
me at first: meh whats the difference?
LLM-based text encoder
me: ok you have my full attention
i like illustrious but gawd i hate prompt fighting
hopefully this one will sort it out or at least lessen the frustration
my only nit pick is lack of base realism (kind of general all in one) support but i guess you get something at the cost of something else...
U actually can get good realism, if u finetune Lora or checkpoint I did train 1 Lora for test and the encoder did good job at caption realism model
@reilgun thats promising to hear
I'm really looking forward to the day it's finished. Compared to version 2, version 3 seems to have more accurate overlapping limbs when generating multi-person images. Is there any plan for Anima to support Chinese natural language? I've tried using Chinese but it doesn't generate the correct content.
I don't think so It would be great if all the models Support All the world languages
It's not a Chinese model. If you want Chinese text, use a Chinese model.
Also, at this moment, Anima does not generate text on images.
creator is great, but the problem of fingers and toes needs to be improved
It seems a matter of gitgudding.
I'm generating them just fine.
In my experience that problem is mostly introduced by loras
I've noticed they can be hit or miss, a lot of my good ones have kinda unusually long fingers too.
Also had a concept where... it kinda understood the idea but it just could not execute it with the hands lol gave me extremely deformed hands trying.
Let's try this bitch out. I'll write back if it's RIP Illustrious moment
Honestly, it's good but not that great. It follows natural language prompts 4/10 times. Simpler stuff is good, especially if you want you to have some text rendering. Groups of people turn out pretty detailed but the models has a hard time following your want.
For corn stuff, it's mostly danbooru tags all over again. And I think IL is much more consistent with it lol.
IL is still better, maybe even much better if you want to do corn stuff knowing danbooru tags.
At the very least, IL is fast af, anima not so much
So We can't use this with Illusitrious Loras ? damn
Illustrious is SDXL which means years old.
Anima is a completely different beast, even though it is only 2B parameters, it uses them incredibly well. It is in some ways better than Qwen for anime.
And compared to SDXL/Illustrious, Anima is a hundred times more detailed.
It's a different model architecture, so no.
@cooperdk yea i like that, the only issue is trying to draw unpopular characters with little to no online images to use s reference, Loras are needed to bypass this issue.
It’s getting better and better.I hope the full version will be released soon.
Nah need more
What scheduler would you recommend using with er_sde?
Made a test with preview1 version Anima CAT. Here is the comment I left on that model's page regarding this. In my opinion "Beta" scheduler works best (currently not available in civitai generator), at least in terms of prompt adherence.
There is an image post attached to the comment, where you can see for yourself.
No idea what i am doing wrong, but me and my friend spent HOURS trying anima but every result (except one) were AWFUL. we used a site for artist styles (anima 2B) and tried a number of different styles that I liked but all results given were like SD 1.5 but with zero negative prompts.
My friend knows what he is doing as he was the one tutoring me through with it but i gave up after around 3 hours of trying and yet people are saying it's the best ever and no one can top it. yet for me i went back to ILL and the very first generated image was great and high quality.
No idea what the magic is about with this Anima but it sure doesn't like me using it.
You definitely did something wrong. And you do not prompt this with booru tags, but in natural language.
@cooperdk wrong, you do both for best effect, first you write in tags, then you write in natural language.
in anima you need @ before artist name for strong effect but to tell the truth, it's not good. Copy paste people prompts (+seed, scheduler and sampler) and see if it works (reproduces) their image. If it does, good, means you were just prompting in bad way (=bad results). If it doesn't = you set up your workflow wrong, drag image from site and drop it into comfyui to copy workflow (from anima prev)
Anima does something better some, worse, certainly its only model that can do something coherent at resolution as low as 320x320, and it can do good quality renders (good enough) at 8 steps. But it's 4 times slower than Illustrious, and unlike chroma/flux, it's not as good with text. But chr/flux are 10 times slower than illustrious so fair game?
I have the same problem this looks like basic SD 1.5 to me too
@AltairTheArc I really like this about Anima because just natural language alone can be kind of a pain in the ass and Danbooru tags are very simple. So it is really the best of both worlds.
Anima quality is 1024x while Illustrious 2.0 is like 1536x.
Theres a whole bunch of other reasons. But basically as of now Anima is no where near as close to Illustrious as most people think.
Its a good model and all, but it fits somewhere between Pony V6 and Illustrious in terms of quality.
@LatteLeopard that's what i see too, illustrious is still the best imo but, with what people are saying anima does well with multiple people in a single image which i try to do sometimes with regional prompter in illustrious and a lot it doesn't go so well.
but other then that I don't see anything so amazing that people are saying about it like "it will change anime forever" kind of vibe people are giving off about it. especially since we're still in the stage where AI cannot identify a lot of characters from anime,games,visual novels and so fourth.
@atmogenic
lmao, yeah your experience is basically what most people feel. They try it, get some mixed results and then go back to ILX or Pony.
Honestly, I dont think it will beat illustrious, and im starting to think that not might be their goal after all.
This model just wants to focus solely anime illustrations and be the best at doing so.
I imagine its final release will compete with Animagine XL
This should put into perspective how intensely the focus is on anime.
Anima is a 2b parameter model thats mostly anime, it does use LAION-POP and deviantart data. LAION is considered cutting edge, but they pruned anything that isnt 2d illustrations.
Also for some reason they used 2MP(1024x) images while ILL uses 4MP(2048x)
Illustrious is around 2b parameters too, but its a powerful mix of e621, danbooru, pixiv, etc.
its based on SDXL, which is reliable, predictable and comes with LAION datasets already trained.
Anima uses a mix of Qwen(for text encoder + VAE) and Cosmos-Predict2,
the base model they trained on was a custom version of Cosmos-Predict2-2B-Text2Image
Actually, anyone can try training additional datasets from huggingface onto it. That would give it more danbooru + e621 data. Also it would increase the quality, because the datasets use 4MP images.
yeah i think but needs months
Thank you so much for your hard work.
The model looks amazing.
Also, would there be an "XL" version of the model? It's only 4GB now, and it seems like there's still a lot of room for growth.
How to do a 2nd pass with this model? I upscaled the image after the first pass using an upscale model, and then i did a low denoise 2nd pass on the upscaled image, but it was completely distorted and pixelated. Will this only be possible in the full version? I saw the hires lora but 2048x2048 is still too small. Sorry but images without upscale + 2nd pass to refine do not look good. A 1st pass image is unusable, too small, too blurry, too pixalated.
Post bugs and bad stuff here https://huggingface.co/circlestone-labs/Anima/discussions
Just want to ask : " the top score is score_9 or score_7 ?" Thank you so much ~
score_9 just like pony^^
just want to say thank you for the new model! It looks amazing and has a lot of improvements. Thank you for all the effort from your team.
May I ask why the model generates a penis in green?
It's a Greenis
Orc dick supremacy, the model is racist
what is dataset cutoff for v3 and this model is smaller and better than sdxl family because of its llm text encoder but it also makes it equally twice as slower then sdxl which is a big let down.
likely dit architecture is the reason
First, thank you very much for the models. I even stopped using IL as my main model type after seeing how powerful Anima can be. But am I the only one getting better details and way better prompt adherence with Anima Preview 2? I tried finetunes, too, and in all cases I'm getting better results with models based on Preview 2 instead of Preview 3.
Hi, is it possible to use "Anima" on AUTOMATIC1111?
If you are referring to the web UI, then yes, I am able to run it using ForgeNeo.
https://civitai.com/images/127669768
As shown in this image, you need to specify the Preset, VAE, and Text Encoder.
Therefore, you must use Forge Neo, which is a fork of sd-webui-forge-classic.
Really looking forward the full model. This really reminds me of my experience with old Dalle, and that could just be nostalgia talking, but I'm really enjoying this one!
Are you sure you didn't use any synthetic data? I feel that several of the generated images have fairly obvious Niji features.
It's what happens if you 3d and 2d data in a model - the average look is the known slop look.
forge doesn't recognize model.
The main branch of Forge hasn't been updated in like two years. Try using Forge Neo; it's regularly updated and has anima support: https://github.com/Haoming02/sd-webui-forge-classic/tree/neo
Just tried. This model do not support artstyle fusion as much as how IL models do. It's gonna largely relies on the artstyle LoRa in the near future
Also the author is somehow planning to delete artist tags for further commercial plans, so guys, don't be overconfident
Through experimentation, I've found out that tagging art styles such as "[@artist1, | @artist2, ]" can merge art styles in a semi-consistent manner compared to the more common "@artist1, @artist2, ". It's worth a try.
As for the "deleting artist tags" thing, got a source I can read? That's genuinely upsetting to hear and it'd be a huge shame if true
@Cazex Sry I was mistaken. It's character names may got removed. Cuz he's actively seeking commercial licenses, so considering midjouney that might happen. But you're right about how it's unlikely to be like that
https://huggingface.co/circlestone-labs/Anima/discussions/37
@EmonDante LMAO you have brain damage
@Cazex yup, use it too. Also helps [@artist1 | ] just to make some styles weaker (some styles just tend to overwrite pretty much anything)
You use a1111 or comfy?
Guys, TDRussell just flew over my house and told me he's going to stop making Anima in 2 more weeks because he's bored. Then he spit in my mouth and told me that no one would believe me.
@EmonDante Ctrl+F there's literally 0 mentions of "name" or "character" in that entire discussion. What are you smoking?
It makes nearly no sense to first train a model on tags such as artist or character names and then remove them... you would be nearly back at square one in the training process.
Trying to use the model and it keeps telling me that it ran out of memory or another weird bug with a lot of words. This is on the example image too. Don't know what I'm doing wrong
why image to image Will turn gray Or add noise.
good when it's good, but very unstable, hands can be decent or SD1.5 tier depending on action and seem to get worse when adding artist tags?, sometimes randomly piss filters all over an image and prompting cool hues turns characters into na'vi, but i'm really excited for future versions
Terrible default art style. 0 art style support. This model is astroturfed.
there is art style support, you just have to add an @ at the beginning of the artist tag.
I DON NOT AGREE !
tfw civit users are too retarded to read documentation
you can train your own art styles bub
Even if you ignore the fact you can just prompt for an artist style, why would you want a model that actively has a strong style bias with no artists? That's how you end up with WAI or Pony...
In my opinion, the effectiveness of artist tag is still quite limited compared to other models. However, its natural language processing capabilities are truly outstanding, and I look forward to the official release.
I’m not sure if it’s just me, but when generating Preview 3 on Forge Neo, it seems best to use LoRAs made after the release of version 3; LoRAs trained on Preview or Preview 2 sometimes cause the output to break on Forge Neo. LoRAs really are crucial.
you dont have to make your message in bold letters, you are not special.
What are the advantages and disadvantages of this model compared to Illustrious?
+ great prompt adherence
+ prompt not limited to booru tags because of natural language
- still on beta phase
- stuck on 1024x1024 pixel
details up x 1000
lo malo no hay muchos modelos y lora por eso seguiré otros
The bad thing is not in Illustrious
After extensive testing 1920x1080 12 steps with your own well-trained lora is all you need for Anima. It just works the same way as the official Turbo lora. Absolutely a waste to run 30steps and you don't need the official highres lora to boost to 2160x2160 square coz the native 1920x1080 limit is much more practical in production stage.
I think this model have been over trained on some characters.
running prompt like this ''masterpiece, best quality, score_7, murakumo_(kancolle) , sailor collar, skin tights, sitting on boy, straddling, 1boy, boy fully visible, '' negative: ''muscle, worst quality, bad quality'',cfg 5, 30 steps, ER SDE beta on forge neo will give you pov, from behind image of that girl, no variety just almost same pose.
dark magician girl have the opposite case, she have many variety image of same prompt I tried.
its giving you exactly what you're describing, which is nothing at all. you prompt 1boy and nothing else - its giving you a trap murakumo straddling the only thing it could with a 1boy prompt - pov
@plan_truster doesn't matter if you add 1girl, 1boy, hetero, same result. I have more prompt if you want to see them. The same problem rarely happens if you try on character like dark magician girl that I mention.
Add pov to the negative and "full body" instead of "boy fully visible" to the positive. And then add an angle. Like your issue is that its biased towards from behind - why dont you prompt something different like from side or straight-on. Or add from behind to the negs? Its not that complicated, yet you - without even trying - race to conclusions about model being overtrained?
@deitychaser my man read my comment, again. I did not even said anything about it being over trained in general. I even used another character like dark magician girl that work perfectly fine from the same model. Other base model like noobai does not have the same problem with murakumo where the same prompt (not the one I mention) give different variant of result of the prompt and that is something i notice and wanted to point it out for improvement.
My guess is that prompt is too vague, add some description to the boy or the scene like:
pos: "score_9, score_8, score_7, masterpiece, best quality, amazing quality, very aesthetic, high resolution, ultra-detailed, absurdres, newest, safe, 1girl, 1boy, murakumo \(kancolle\), kantai collection, sailor collar, skin tights, sitting on person, sitting, on chair, straddling, full body, the boy is sitting on the chair"
neg: "bad quality, worst quality, low quality, score_1, score_2, score_3, lowres, distorted, signature, watermark, patreon logo, artist name, poor lighting, worst detail, multiple views, logo, patreon username, web address, dated, bad hands, extra arms, extra legs, amputee, mutation, extra digits"
@FinisherStrike Look like removing masterpiece and best quality and adding score_9, score_8, score_7, gave better result. I tried that on fate grand order character reines el-melloi archisorte. it gave better result.
Any recommendations for imitating poses something similar to ControlNet(Forge NEO)?
I made a bit of experiments:
Using the same parameters from TensorArt on my local pc results in two wildly different images. (I'm using Forge Neo).
I assumed the possible culprit was the VAE, so I ran some tests, but besides 'qwen_image_vae' on my local computer, all other VAEs result in black or grey images.
Using 'qwen_image_vae' produces different images compared to what I was getting on TensorArt, so I tested again, and I ran using the same seed and parameters, but different VAEs on TensorArt, and I got identical images each time.
Therefore, TensorArt lies in what parameters it is using with this model and does not tell you what it is. Kinda bummed because it produces better result that what I'm getting on my pc.
very GOOD model!!!!!
it can understand nature language.
thanks a lot!!!!
BTW,
is it possible to train rare prompt, like bodypaint/scat?
(AI translation)
Compared to my go-to model WAI-illustrious-SDXL, Anima has clear advantages in natural language understanding and artist style replication.
However, there's one major issue with Anima that I find quite regrettable, though I'm still looking forward to future versions.
The issue is a bit strange to describe.
WAI seems to have much stronger style mixing capabilities. I can easily blend the styles of different artists, fine-tune the current style, or create something between two styles.
Anima can mix styles too, but the result feels strange. It seems like it's not just mixing visual effects, but also associating semantics.
Here's an example using akai sashimi and modare – both artists produce work that is simultaneously realistic and not realistic.
Akai sashimi's work is very realistic in subject matter choice.
Modare's work is very realistic in texture rendering.
But both share a common trait: they generally don't draw nostrils.
When I mix akai sashimi and modare's styles in WAI, the result is very intuitive. I don't feel any issues at all.
When I do the same mix in Anima, the result is… odd. It's like it kind of looks like them, but also kind of doesn't. It feels like someone who has never actually seen these artists' work is trying to guess and draw based on someone else's verbal description. And the probability of characters having nostrils is strangely high – something that simply doesn't happen in WAI at all.
(AI translation) Yes, I am experiencing this issue as well. And, if you pay attention to the feedback from the community, you'll see that you're not alone in this matter. The problem I encountered was not only that the artist's style could not be blended together, but also that some artists (such as @michiking) still had subtle differences in the AI-generated image style even when used alone.
Artist mixing is a consequence of CLIP based models like stable diffusion XL. (Pony, illustrious and all of it's mixes) Every model after stable diffusion abandoned CLIP in favor of LLMs for their understanding of context and precision. Regardless of which newer model you use, you won't get artist mixing like SDXL again. Your best bet is to use something like Prompt Editing, it's not the same but it's better than what you get with just listing artist. I wrote an article on the subject here.
Yes, artist mixing is the legendary "bug becoming feature."
Deviantart dataset seems to have ruined the model. If you preview steps 2-4 there is a clear deviantart logo right in the middle of the image. Artist and style tags are also significantly less impactful compared to p2.
how to emphasize the prompt to make it weak or strong? like I want to make some artist tag only 0.5 strength
i use (@artist:0.5) and it seems to work
For some reason, v3 it really doesn't play nice on classic. It outputs things, sure, but does not output things as accurately as in Comfy.
Any recommended Style Lora guides? I see this account posted a "Greg Rutkowski Style" guide. Is that the best settings?
i think that was posted for an example lora with training dataset and settings. as op is owner of https://github.com/tdrussell/diffusion-pipe training toolchain.
Very good model, I have high hopes for how it will shape out to be, but preview 3 is already very competent. Good work guys and gals!
I've strangely found this model generating significantly better and more consistent results at specifically resolution 1248x1248.
Can anyone else check this out and let me know how it goes?
the new model in general generates higher quality better, i downloaded the hi-res fix to do 1080p and 1440p wallpapers and i did some 1080p wallpapers first and it was perfect, just to see i had a lora for 3dcgi enabled for illustrious and not the anima hires fix enabled lol, can do around 2.5MP decently in my tests
This model surpasses existing SDXL-based models such as Illustrious and NoobAI in several areas, including natural language prompt understanding, English text generation, separation of multiple character features, and fidelity in character and style reproduction.
Thank you for sharing such a great model!
I’m looking forward to the release of the full official version.
Hoping for a Anima-preview lora online training, it would be really cool 👍 this model is too peak
the creator want to sell license, so it will take a while for online lora training. Honestly, this model made me to train a lora on my 5 year old pc.
I haven't been able to get good results yet with lora but I've been using someone else's preset since I'm unfamiliar with this model
In the next update, I'd like to see more characters (specifically from One Piece) and improvements to the art style. Additionally, the '@oda eiichirou' tag requires adjustment: the style strength is currently too low, and there is a persistent white border that cannot be removed even with negative prompts.
If you look at the images on danbooru tagged with "oda eiichirou" there are a lot with white borders, so there is little you can do about it I guess. You can try using @oda eiichirou \(style\) to get a similar style .
@RisingV I'm still jealous of the capabilities of NovelAI. They are far more advanced in datasets of art styles and updated characters. It's just that it is closed-source and you have to pay to use it...
@monicalucci Yeah, open weight models often have fewer capabilities...
But that's what LoRAs are for. If they could do every character out of the box, LoRA creators just as myself would have little to do ;)
I like that this model is not only anime style but also have decent cartoon art style in it without writing any artist.
it can do more of a western illustration style too which I enjoy, 2d models often have that ugly AI slop look to them, anime models make the eyes too big. imo it can be difficult to get that kind of style repeatedly.
@minthe the creator said he used 800k images from other sources. I assume he was talking about cartoons/western style
In Forge-Neo, I'm running into a black image and a message that says, "Encountered NaN in Latent".
I'm using the expected text encoder and VAE.
I couldn't find anything, looking around, does anyone have any leads?
I'm not entirely sure but I think it might relate to the sampler and scheduler combo. If I do Euler A and Automatic it gives me a black image with the same NaN error, but if I do Euler A and Normal or Simple then it works fine, same prompt and everything else.
@bloodmetaleyes6727
That seems to be exactly it. Huh. Yeah only Normal and Simple work and all other produce a black image.
Thanks for the response dude, you're a saint!
I hope you guys continue to support natural language prompts in your models!! it does great job!
The model is very promising, but it's clearly a bit rough around the edges. We're all delighted with the consistency of the scenes and characters out of the box, but there are a couple of nuances worth noting:
1. The dataset is obviously poor, making it very difficult for the model to grasp many concepts without detailed descriptions, but I understand that this is a downside of the preview version.
2. The very limited context means the model's attention starts to blur even with four characters with different eye colors, hair colors, and facial expressions. This is especially true with just portrait-style focus on the faces, without complex poses. Experiment and you'll notice that character 1 is more or less fine, but the emotions can blend with those of character 2. Characters 3 and 4 simply average out, and sometimes the 3rd and 4th characters swap positions, ignoring the prompt. These are all clear signs of a lack of attention on the model's part, even in such a relatively simple scene for a transform base.
3. I don't know why, but the model reacts very strongly to specific triggers. Just one short prompt can radically change the model's understanding of the scene, while the model might simply ignore a detailed prompt.
But despite all the criticism, the guys did a fantastic job. The basic goal of bringing consistency capabilities into the hands of the public is invaluable. And as I myself have noted, most of the model's drawbacks are more related to optimization decisions or the experimental nature of the model previews. I wouldn't call Anima a full-fledged "Illustriuos Killer" yet, but the model already offers capabilities that no anime SD model has offered before.
in 2. I think it more to do with the text encoder, while it is way better than sdxl text encoder, it struggle also with short comic page so it is not surprising it struggle with multiple characters.
I think it more to do with the text encoder, while it is way better than sdxl text encoder, it struggle also with short comic page so it is not surprising it struggle with multiple characters.
@Dewal76 Yes, that's right, Qwen 0.6B clearly has too little context, which is why the model starts to generalize and forget details when describing in detail. It would seem that we've finally gained the ability to create multi-character scenes out of the box, but the limited context spoils a lot. However, if we move to a larger Qwen model, Anima risks becoming too heavy and less accessible to the masses.
Not to dampler your expectations for the full release, but what you are asking for is probably not feasable for such a small model. It's a locally deployable model with a tolerable generation time on average hardware. But that doesn't mean that regional prompting and other nifty techniques for multi character scenes have become obsolet. Use them.
I would like to point out that as per the HuggingFace discussion on "prompt weight adjustment", this model does in fact support messing with the weight of input text.
It is just that instead of "1girl, (gigantic breasts:1.5)," to get a ridiculous result in SDXL, you instead have to specify "1girl, (gigantic breasts:4.5),"
different things are more sensitive than others (e.g. '@artist' tags), you can even play with entire phrases like "(cute goblin girl surfing on a giant leaf in a cloudy landscape:1.15)" to see what funky results it summons
and yes, this also goes for NegPiP in comfyui-ppm, where you can do hilarious things like ", (worst quality:-7.5)"
turns out this is just as much FUN as SD1.5, except now we have several really strong DMD LoRAs that can be adjusted to make the composition coherent, as well as a CFG distillation, surprisingly good 'concept isolation', and a model that is not prone to collapsing when you push it
so, really, it's even more fun, ugh.. so many sillies to try
Well done for duplicating this lifehack here. I recently read about it myself. The ability to adjust the promt with weights really does greatly increase the usability of the model, especially considering that the model is extremely uneven in its perception of token weight.
So is this model so good already? stable? I tried it and it wasn't good because I didn't apply what you mentioned, unsatble models like NoobAI Chenkin etc models or Newbie and Lumina models were not useful for me, but it seems that with what you said Anima can be controlled better, in this case it would be really helpful.
@Suomsoh This model is shockingly good, as we get to toy with schizo SD1.5 style prompts, except now with a model that doesn't collapse immediately, has surprisingly good composition and concept isolation, and doesn't have a 75 token limit before quality of outputs goes down tremendously.
https://civitai.red/images/128869511
Here's an example image with a very distorted prompt, yet it could accomplish the challenge I set for it
furthermore;
https://www.dropbox.com/scl/fi/ewxt950rxoh96xiezlnun/ComfyUI_anima3-stuffs-minimal.7z?rlkey=5d3jm165295pwnzirp71vl4yf&st=l2lqtkxf&dl=1
Here's the necessary comfyui-ppm, LoRAs, and example images of a working wacky-prompt workflow in one package
@Suomsoh and yes, on its lonesome without Wacky Setups(tm), it's rather hard to get a convincing output from the model when you stray from "Your Average Waifu Printer" territory
@VeerGeer Thanks man, I replied a week later haha but I'm just trying Anima again so yeah it's good, hope it gets proper controlnet support, lora "controlnets" doesn't work too good haha, this to add to a tile workflow so that it works better with a decent denoise setting
@Suomsoh With how "fun" the model is, it's almost guaranteed that it'll have a good lineage.
I have grown quite tired of all these "modern" models with their "insert LLM prompt or perish" requirements, and it's rather humourous to see people stumbling into issues with "natural language prompts" when said 'natural language' is in fact, not written by them.
@VeerGeer Hahahaha yeah and for anime models tag prompts are needed the most in my opinion, like why in the world would an anime model not need tags? even a realistic model needs them in my opinion, if possible just having a text encoder for tags and another for natural language to merge them both correctly would be cool, if not then I prefer tags
I tried Z-Anime but it was bad because of the natural language only prompts how funny to train a model that was made using a Danbooru dataset of images uploaded with tags to then make it only natural language hahaha what a waste of categorization/organization/information etc etc and indeed, like you said, lots of people doesn't write the natural language prompts yet they say that tags are worse when in fact they aren't, at least to me, tags are easier to use than natural language only prompts
And not like I just use AI for nsfw content, but Z-Anime used the dataset of NoobAI that was uncensored and then they censored most of it so they made it the most sfw possible, so to me this directly is a model I won't be using now that there's Anima all uncensored and knowing that the ones behind Z-Anime want the model censored, like why haha so to have reputation of being family friendly or something like being on good terms with those that say that the people who use open source models are "gooners"? ok whatever, I hope there are more Anima uncensored models in the future or similar ones, but if not then Anima will be the last model I will use for Anime haha I don't care about those new lots of gbs models that output mostly censored trash or looks like plastic/shiny skin/double chin/wierd AI aesthethics/blurry backgrounds models
This is a version, but in illusion.
Is this somehow working in Automatic1111?
Imagine what version 4 can do
I hope this model can achieve similar results to NovelAI.
This requires deeper training, as well as training on the e621 dataset...
which most likely won't happen.
@socksosock Yes, I totally get it. NovelAI is a proprietary model with a much larger dataset than Anima, and it can understand way more tags and captions. That's a really hard gap to close.
where do i put this in comfyui? in checkpoints it gives me error
diffusion models and load diffusion model node
is there a list of characters, and concept tags this recognizes ? Also please refine poses for future versions. Its obsessed with ass shots.
Maybe put "safe" in the prompt to avoid provocative poses?
What does the dataset look like for this? I'm uncertain how to tag/caption when training LoRA for this since it uses both and tags don't seem to like the periods used in captions. At least not in edit software.
Their official hugging face page have some details about the tags it was trained on: https://huggingface.co/circlestone-labs/Anima
Mainly trained on anime images with additional datasets of non-anime artistic images (ye-pop + deviantart). For captioning especially refer to this comment by tdrussel (the person behind anima) on the anima huggingface page: https://huggingface.co/circlestone-labs/Anima/discussions/9#69812bd9511f2d67952084ae
Using tags only in lora training seems to work fine in my experience.
@RisingV Thanks! How many steps have you been doing in training for Anima by the way? For awhile, rule of thumb was 2000 steps but I'm getting really mixed info with optimizers and learning rates. For instance, a few people recommended I use 0.0002 or 0.0001 with Adam but that hasn't worked out well at all with that total of steps.
@minthe For a simple character LoRA I'd recommend lr 1e-4 (0.0001) with 2000-3000 steps using AdamW8bit. May work with less steps and higher learning rate though.
I have added the total step count to the details of my models, but unfortunately they are not displayed anymore.^^
Here the step counts and learing rates I have used for my character LoRas :
https://civitai.red/models/2586297/bibi-blocksberg-little-witch-anima: 3000 (lr 1e-4)
https://civitai.red/models/2513900/nami-the-thiefing-cat-straw-hat-pirates-romance-dawn-arc-one-piece-animail: 3000 (lr 1e-4) + 900 (lr 5e-5) (for two outfits)
https://civitai.red/models/2405649/nami-the-thiefing-cat-romance-dawn-orange-town-syrup-village-arc-one-piece-animail: preview2 version: 1470 (lr 1e-4), preview1 version: 1260 (lr 1e-4)
I guess you have to try a bit which lr/steps work best for your dataset. If you have a very diverse dataset in terms of style you might want to use more steps than with single style (e.g. screencaps). I usually do not use large datasets for characters (about 30 images for single outfit loras).
@RisingV Thanks for sharing! I've used that lr and 0.0002 for style lora but they haven't come even close till the last epoch. I've only ever managed to get 0.0005 to work with my dataset, might just have to go back to doing that even if it overfits a little lol
@minthe For style loras I'd recommend even more steps with a lower lr. On the anima huggingface page it is recommended to start with lr 2e-5 and go up or down from that value. That is maybe too low, but if you use a learning rate that is too high details might not be learned regardless of how long you train. Then again I don't have that much experience with style loras.
However I have trained a style lora using a limited dataset of 65 images with lr 5e-5 and 6500 steps: https://civitai.red/models/1890171?modelVersionId=2847255
It looks pretty close to the images in the dataset I think.
Also be aware that images generated with style loras can look pretty different depending on which scheduler and sampler you are using for generation.
There is a style LoRA published by tdrussel with the training settings in the model description, if you didn't know:
https://civitai.red/models/2536147?modelVersionId=2850290
@RisingV I've gotten so much conflicting feedback, it gets confusing lol. Someone told me 2e-5 with CAME and they got what they wanted 1000 steps in, I had over 4000 steps and still nothing, another guy says 3000 is too many, while one says too little.
Thanks for letting me know about the tdrussel LoRA, I didn't know. I think that'll be the most helpful to take a peek have, though I assume he uses a lot of images for it.
@minthe Yeah, even though anima preview is out a few months now, I guess we are still pretty much in experimental stage regarding LoRA training. If you used SD models before, the switch from SD 1.5 to SDXL or within SDXL models from Pony to illustrious was probably easier since they are all based on the same model architecture (more or less).
For the first character LoRA I published for anima I actually trained 8 versions with different settings to figure out which works best for the dataset.
However after training more than 100 LoRAs with illustrious I am still experimenting with different settings/training strategies for that to get better results.
Results also depend on other settings than lr and steps, but these are the two parameters I am usually messing with.
Maybe I should share the full settings I am using in an article (not that they are really good, but they work for me).
Haven't tried CAME (ever) so I can't say anything about that. Have you been training the text encoder (not the adapter)? Only training the transformer might be slowing down learning a lot.
Regarding the Greg Rutkowski style LoRA by tdrussel, you can download the training data used for that one, it's listed under optional files in the model card. 153 images were used (that's a normal amount for a style lora I guess) so the total training steps were 18.360. Probably you don't need that much steps if you use a smaller network dim, especially for higher lr.
Another thing I noticed for styles is using "@" in front of a tag/trigger word will amplify the style effect of said tag during generation. That does not only work for artist styles the model already knows, it also kinda works for series styles, e.g. "@one piece". The same applies for trigger words of style loras, on the Greg Rutkowski lora page it is said to put "@greg rutkowski" in front of your prompt to get the style working. Here are two images with the style lora I trained, in the first one I used just the style trigger word and in the second I put "@" in front of the style trigger word while keeping all settings and the rest of the prompt the same:
https://image-b2.civitai.com/file/civitai-media-cache/a1c3c43e-fabc-46a7-854c-1c30ae539756/1200x%3Cauto%3E_su
https://image-b2.civitai.com/file/civitai-media-cache/6e518b78-3388-451a-b0cc-cf8fc68b95cf/1200x%3Cauto%3E_su
As you can see there is a visible difference between the two images and the second one looks much closer to the trained style.
I hope this helps you somewhat.
@RisingV I heard text encoder is not very useful with Anima so I kept it turned it off, so only training UNET. But CAME apparently makes training slower, I heard it was based on Adafactor. Although no one seems to use Adafactor anymore!
I went back and used the @ on my trigger word, as well as other tags, the style actually did appear as early as epoch 2 but it was still... kinda wonky. I might give up on that one in favor of trying anime, though I had a similar struggle getting it to understand!
Yes this has all been very helpful, so thanks for taking the time to explain it to me^^
@minthe You should really try training the text encoder (qwen llm) as it makes getting the concept/style easier you want to train, just be sure not to train the llm_adapter as this might damage the outcome.
I have been training the te by default for my illustrious loras as well, so I can't really tell how much it's influencing the outcome. However I tried to forcefully merge one of my style loras into anima-preview1 checkpoint without really knowing what I was doing. The result was only the DIT (transformer similar to UNET) part was merged and the text encoder part was discarded. The result kinda looked like the style, but there were certain aspects missing compared to using the style lora with the original preview1 checkpoint.
I am not completely sure, but I think training the text encoder makes the model relearn the meaning of the tags you are using in the caption of your dataset and not just the aesthetics of the images, if that makes any sense :D
I have only ever used prodigy when I started doing LoRAs and then switched to AdamW8bit and stuck to it.^^
No problem, if you have another question, just ask, though I am not sure if I can give good answer. ;)
Can you upload them to pixai please? Sooner or later someone will and I think its better if the creator gets the credit
pixai is trash They only have SDXL models
You can't just upload a dit model to pixai- it needs to be setup correctly on the backend by the devs of the website (pair the correct text encoder and vae with the correct diffusion model). So either the dev team wants to host the model in which case they will take care of it or they'll keep pushing their own dit anime models that they developed.
@Iwsnsiwis282 What the heck, PixAI has their own DiT models now, Tsubaki series
@deitychaser This is true, either PixAI and Circlestones Lab has agreement to host Anima model in PixAI. But I don't think they are going to host other DiT models now since PixAI has created their own anime DiT models.
@springmushroom_86 They have to buy the license like tensor art and this website
How can I know which version of name tag of artist is correct? I wanted to try this model with Miro Haverinen art but he also have Happy_paintings nickname. Can you clarify this for me, where did you look for the artists name?
I don't think this artist's artworks have been posted in any booru website. You may need to find stye lora of this artist first. If you want to explore more danbooru artist tags in Anima, you can check this website: https://thetacursed.github.io/Anima-Style-Explorer/index.html
It will work for most artist styles with at least 150 images uploaded on danbooru. That said some artists may require using their gelbooru tag instead of the danbooru tag (in case they differ).
I love that this model has normal art It gives him a soul Art doesn't always have to be high quality
does it do nsfw?
Yes. Just browse the gallery below for more examples (with PG/PG-13 disabled).
@RisingV ok thanks
i also wonder if it's backwards compatible with illustrious LORAs
@oh_well Anima won't be compatible with Illustrious loras due to different architecture and training.
@springmushroom_86 oh thanks (2)
1 month ago I wonder what they will add this time
I just get "error in loading state_dict" for the text encoder with a long list of size mismatch errors.
Make sure you are using the Qwen clip and vae linked in description. In case your UI doesn't auto-use them, you may need to point to them manually.
update your comfyui
Holy. This is some dark magic shit we dared not to think possible two years back. Crazy natural language adherence and easily runnable on potato PCs.
if its already this flexible on just the preview version, i cannot wait to see what v1 can do 👀
Well, flexibility probably will not improve with further training, but you can expect better highres and maybe some additional style/character knowledge.
@RisingV To be honest, this model has some issues—its composition is too rigid, and vague prompts almost always yield highly similar results, which lacks a bit of creativity. However, it’s a great choice for redrawing models, as the details are very clear even at low resolutions.
@2253271156 I guess it's not as "creative" as illustrious or other sdxl based models and you have to be more precise in your prompt regarding what you want. That's probably because of the different architecture (DIT instead of U-Net and LLM text encoder instead of CLIP), but their are some efforts to work around that, e.g. https://github.com/Anzhc/Anima-Mod-Guidance-ComfyUI-Node.
To be honest, this model has some issues. Its composition is too rigid, and vague prompts almost always yield highly similar results, which lacks a bit of creativity. However, it’s a great choice for redrawing models, as the details are very clear even at low resolutions.
the problem comes from the tags, the ''masterpiece, score_9, highres, newest, amazing quality, '' etc narrow down the result so much that the model produce same images. it basically have the same problem ponyxl had. use instead fine tuned model with styles baked in without using any of the quality tags.
You can use significantly less steps on preview3 than on preview2 imo. I get on an old prompt for p2 same diversity when i go down from previously used 30 stepst o 15-18 steps.
help, forge showing ValueError: Failed to recognize model type! found out that forge is unable to recognize the file, tried changing filename and adding a .yaml file in models for forceful recognition but still not working!!!
I have to say, this model is excellent, with an unexpectedly high ceiling. The issue lies in its handling of multiple fingers and fine details. As you can see, it tends to produce extra fingers during image generation. Of course, I believe this should be resolved after the official version is released.
Use "four fingers" and "six fingers" in negatives. It should help.
it's a rite of passage for base models to fuck up hands
Sage attention 3 make the model ouput real people instead of Anime. -_-
What's the best way to prompt for on-screen text, especially speech bubbles? I mostly get broken output.
I use something like:
speech bubble, text: "a few words max"
You can also do stuff like:
there is a speech bubble above her head with a powerful font saying "BLAH BLAH BLAH"
for more control
There's also a lora called textify which can help a lot sometimes
@9643 Thank you for the tip with the textify Lora. Only tried it a handful of times just now, but the early results are great!
What's the clip skip number to use? Still 2?
Practically, Clip Skip never mattered after SD1
Any method to reduce limb distortion? It seems that with every other generated image that the limbs come out lanky or bloated even though the prompt doesn't change.
Can I achieve that "only change clothes, strictly doesn't change character and poses" with this?
Can I achieve that "only change clothes, strictly doesn't change character and poses" with this?
yeah i dunno what i'm doing wrong, followed the base model tutorial, but i don't get any image to look decent, it's all weird jumbled mess, or like the results aren't anime at all... tried using the tutorial prompts or my own prompt and same stuff... dunno what to do... I'm using forge neo
you didnt dig deeper yet. you must be a gen z or alpha
@jancok you could maybe offer some advice of being condescending... oh wait that's right i'm on the internet, silly me
Here's a checklist of all you should theoretically need:
vae: qwen_image_vae
clip: qwen_3_06b_base
scheduler: sgm_uniform
sampler: er_sde
And the standard 4.5-6.5 cfg with 20-40 steps
If you've got that whole list checked and it's not working, then I have no clue I'm afraid
@9643 my previous setting was everything you said, except the scheduler, and i doubt this one will drastically change the results, but i'll try later, thanks anyway!
Make sure you are loading it as a diffusion model, not a checkpoint.
Happy tensor.art got anima training now i can make anything i want these people only posting nsfw models 🥀
So, how much of the Gelbooru dataset is integrated on Anima?
Does it work in Forge Neo?
Classic yes, but only certain forks, Neo I AM NOT SURE. Reforge no.
Classic i ONLY know of because i checked the repos.
To be QUITE fair: you're better off with ComfyUI - and this PAINS ME TO SAY this as i am a die hard Forge/A111 user but the problem is: Gradio is VERY difficult.
I am trying to develop my own wrapper UI for comfyui just like some of my peers, to make it easier on people - but i'm not gonna lie: NEO drove me away from Gradio based apps.
yes works perfectly on forge neo.
dont know which version you got, but you might have to update forge neo
@ktiseos_nyx Sadly i agree. Lot of folks recommend comfyui but the thing is, i just can't get used to it and it complicated to use. I would like to see your own UI for comfy.
@ModeComprehensive506822 Great, now that i update Neo, there seems to be option for Anima. Let's see if the checkpoint model itself works.
@chrisss1 nice everything should work now. thats the same way i use it
@chrisss1 I'm working on it, Ktiseos Nyx Trainer will become an Ecosystem with metadata reading and ComfyUI with an actual working react based UI system - im HOPING to have some of this started this week - the trainer works fine for the most part, and i got it to train Anima loras.
I don't want to link to it in here because that's cross advertising and unfair on Circlestone labs.
But i did make an article about it yesterday
I'm having some serious trouble with this model being very... static. If you use a prompt that's more detailed, it cuts the creativity of the model IMMENSELY; the more tags/tokens your prompt is, the less variation and creativity it gives in return--it's almost always just slightly altered versions of the same image structure. I often have to check the seed, to make sure I didn't use the exact same one by mistake.
I'm assuming because this isn't a 'full' model and there just isn't a large enough training pool to draw from because it's a 'preview', at least that's what I believe... but it's becoming a little frustrating having all my prompts having almost no creativity or variation behind them--something that NAI had an overflowing amount of, in contrast.
Any ideas/tips to improve the creativity of the model, beyond just lowering the CFG?
If you use ComfyUI, you can try using different or empty prompt for the first step, then make the rest of the image with your actual prompt. It was used for Z-image Turbo for the same reason. I'm not sure how it affects prompt following, but it adds more randomness.
First KSampler Advanced: add_noise: enabled, start_at_step: 0, end_at_step: 1.
return_with_leftover_noise: enable.
connect the output to the 2nd sampler.
Second KSampler Advanced: add_noise: disable, start_at_step: 1, end_at_step: 10000. return_with_leftover_noise: disable.
steps, cfg, sampler are the same
@NNAI I am using Forge Neo, though I appreciate the tips, thanks!
I guess this is what you get with a DIT/LLM model: more precision but less creativity. You can try CLIP (with modulation guidance) as text encoder instead of qwen llm, at least I think it was implemented for forge neo:
https://github.com/Haoming02/sd-webui-forge-classic/pull/843
https://www.reddit.com/r/StableDiffusion/s/zwZ3Nb3JtG
@RisingV After doing further experimenting/testing, I'm now 99% sure it's just the model itself. There are certain tags that will just LOCK the model on a near identical image structure, repeatedly, no matter what seed is used. I used the CLIP L for a bit as you suggested, and it seemed to have extremely weak prompt adherence in comparison. I guess all we can really do for now is wait for a full release. ¯\_(ツ)_/¯
Mh, yeah can be the model itself. Usually I do not use prompts with a large number of tags, so I can't really say I've encountered the same issue.
I do not think creativity is getting better with more training though. I expect better highres generation with the full model and maybe some more character/style/concept knowledge.
Oh and there is something I have encountered during LoRA training with anima. It seems to pick up details better than illustrious. So maybe it guides the generation to a specific image in the dataset for lengthy prompts? I am not entirely sure how anima was trained regarding tag order, tdrussel made a comment about it here: https://huggingface.co/circlestone-labs/Anima/discussions/9#69812bd9511f2d67952084ae
If it wasn't trained with caption shuffle, i.e. shuffle the tag order for different epochs, the output could depend on the order of tags in the prompt. Have you tried to mix up the tag order (without touching the recommended order)?
@NNAI I think the same or similar method was implemented into Forge Neo by @HaomingGaming for Z-image: https://github.com/Haoming02/sd-webui-forge-classic/discussions/421
But for some reason I can't find it in the settings. Maybe it was removed in a later commit?
@HaomingGaming Oh, did not know you moved it to an extension, thanks for clearing this up! Doing a quick test with it does not seem to do that much for me in terms of improving the output variance across different seeds with base anima-preview3, but maybe you need to play around with the settings for that. Have you tried this @Terrordacturne?
Ok so up until learning how to train on Anima with my own training system (Sorry Civit, I had to graduate at some stage lol) ... TODAY is the first day EVER i've used Anima.
Dude.
Like I love pony and illustrious but let me be that idiot millenial:
"ANIMA!!!! movie dude voice IT REALLY WHIPS THE TENSORS ASS" (instead of the winamp.. the llama's ass.. yea ok i'm gonna walk off now)
As much as my current loras are a LITTLE janky that i just started - I'm REALLY impressed, even WITHOUT the turbo lora and my style lora.
Dunno if this interests anyone but since yesterday an outfit i prompted glitches out (in this model as well as in WAI-ANIMA). something in these prompts makes it so as if the picture was crudely drawn and then colored. If i use a different outfit, i don't have the problems. would be nice if someone could test it, because i don't find my error.
she is wearing body jewelry in the shape of a halterless one-piece swimsuit. (the suit has gold chains that follow the curves of her body. Over her nipples are a small round shape aqua blue gems inside a gold frame that barely covers her nipples and have an areola slip. over her pussy is a heart shape cut aqua blue gem that hangs loosely in front of her pussy. a gold frame connects it to the chains. the gem only covers her pussy from the front). the rest is either exposed or has delicate gold chains running over it. the whole body jewelry is designed and interwoven and connected with each other. the whole outfit is connected and held up by a delicate gold choker.
Try the anima turbo Lora. At least on civit's generator the dress is working well with the Lora on
Update: make sure to also prompt the head, face, etc because it's focuses so much on the dress the rest won't generate
@rotechacon667 problem was an update inside forge neo that broke how lengthy prompts are interpreted. but thx anyway^^
It was surprising to see how well a 8Mb LORA works on this model.
Yeah, I use only half the network dim for character LoRAs I used for illustrious (illu rank 16 is about 100 MB and anima rank 8 is 34 MB).
After training a few Loras using Anima I'm overall pretty impressed but I'm finding it harder to train than illustrious (funny considering how people report having an easier time with Anima), maybe I just need more practice but overall the results are usually better on anima.
I'm also impressed with how well it learns art styles, it's the one thing that I struggle with Illustrious that trains almost flawless in Animal.
Can't wait for anima preview 4 and the actual release
any tips on that? been trying to train it on styles. what steps do you train for. and for the dataset do you use images with only single character or is multiple characters fine?
@dirtydan3213s I usually train 550 steps, Not sure if the amount of characters really matter for a style lora, just make sure the style is consistent. Also, not sure if it's just a placebo effect but add @ before the trigger word. for example: "@artistname"
@dirtydan3213s For styles you usually want have longer training times (more steps) with lower learning rate. Here is an example style lora with settings trained by the anima founder tdrussel:
https://civitai.red/models/2536147/greg-rutkowski-style-anima
It was trained for over 13.000 steps. You probably do not need that, if you use a higher learning rate. I guess using 3000 steps with AdamW8bit optimizer and learning rate 1e-4 is a good starting point. If you use a learning rate too high, details may not bake in. For a style lora you usually want a dataset with as many different characters as possible or the characters generated with that style lora will tend to look like a character from the dataset (so more characters means more flexibility). That being said, I trained a style lora with a dataset dominated by one character and it seems to work fine: https://civitai.red/models/1890171/kyhu-artist-style-k-y-h-u-old-pen-name-of-iafhy-animailzit
glad to know I'm not alone! Training on anima has been a huge pain in the butt lol I've had to get out of my comfort zone and try new methods, especially since I refuse to just use the backend and rely on a frontend still
everyone has such different ways of doing things that is is very difficult to figure out what method actually works for you.
@RisingV thx for the reply! i actually didnt know high learning rate was bad i do actualy train averagely 11k-13k steps. but i do use pridogy with 1 learning rate. so i dont know if thats bad or doesnt matter since prodigy is a special case. or do i just use adam normally since i bet there isnt much of a difference between these things
@dirtydan3213s If you use prodigy a learning rate of 1 or close to 1 is recommended, since it's an adaptive optimizer and regulates the actual learning rate by itself. I think I haven't used prodigy for any style, so I can't say anything about the step count. But since the learning rate is regulated more steps will not necessarily mean better style representation. There are other parameters than the learning rate you may have to tune if using prodigy though.
@RisingV personally, I never had much success with Prodigy for some reason
Well, when I started making LoRAs I read this guide by @Valstrix and the recommended "easiest" optimizer was prodigy, so I was using that and the recommended settings for it to make some character LoRAs and I think most of them were ok. But then I switched to AdamW8bit and never looked back....
I guess you still need the right parameters for it.
@RisingV that guide is pure gold. When I'm stuck with a Lora I always go back and read it.
You're right, it is. It covers everything you need for making SDXL anime loras. I have also revisited it a lot. Though not so much the past months. For illustrious I do not really need it anymore and it does not cover anima or other dit models. And non-standard loras are only touched slightly (probably bc nobody is doing them).
@RisingV I'm sure Anima guides might start coming any moment.
I’ve been running into a persistent issue when using this model: arms and hands almost always turns out distorted or mangled. I’ve tried various LoRAs to fix this, but unfortunately, none of them seem to help.
I’m not sure if I’m missing something? :(
Do you have an example image with prompt? Otherwise it's hard to tell. Should work without using a LoRA to fix it. Have you been using the generation settings and prompting advice given in the model description?
show us your settings, like what sampler, shift, scheduler, etc. in forge neo latest update you have a option called shift. increasing it from 3 to 6 maybe will help. I had weird artifact of cars inside. increasing the shift reduced many of the artifact i had in my images.
Does anyone know how to merge ANIMA LORA with ANIMA LORA, or ANIMA LORA with ANIMA Checkpoint?
I'm trying to do that right now, and I'm lost. Did you find out how?
For ComfyUI
You load the model and LoRA like in any workflows, set LoRA strength, then add a 'ModelSave' node.
Merge into Model:
'Load Diffusion Model' -> 'Load LoRA' -> [optional: chain more 'Load LoRA'] -> 'ModelSave'
Merge LoRAs:
'Load Diffusion Model' -> 'Load LoRA' -> [optional: chain more 'Load LoRA'] -> 'ModelMergeSubstract' (model1: model with LoRAs loaded; model2: original model, connect directly to the 'Load Diffusion Model') -> 'Extract and Save Lora'
lower rank gives smaller files, but may reduce LoRA strength.
if you have KJNodes extension (has more options for extract):
'Load Diffusion Model' -> 'Load LoRA' -> [optional: chain more 'Load LoRA'] -> 'Model Save KJ'
'Load Diffusion Model' -> 'Load LoRA' -> [optional: chain more 'Load LoRA'] -> LoraExtractKJ (finetuned: model with the LoRAs; original: model directly from the Load node)
this I merged and works fine, the size is only 7.7 Mb.
https://civitai.com/models/2498745/nn-semi-realistic-anima?modelVersionId=2930490
this is a merge of the other 2 versions, both at 0.7 strength using LoraExtractKJ.
lora_type: adaptive_energy
algorithm: svd_linalg
adaptive_param: 0.3
@NNAI THKS, now come to downdoad kj node and trying
Another month passed preview 4 plz🥺🥺🥺
This is just speculation, but I think he's planning to keep training until the model is complete. So there won't be anymore preview models, we will have P3 until the model is fully trained. I speculate we will be waiting an additional month or two for the final release. This would explain why he chose to name P3 a "base" model.
@Fish788 tbh I agree with u as they called it "base". But as long as it's preview loras/controlnets ain't committed awww
preview 4 or Final version? They did a lot
Does anyone see a future for this model? If so, what does it look like?
pretty good.
I think this is the SDXL 2.0 everyone’s been waiting for—the anime-specialized version
@CipherVisage But what can it offer that Ill can't? I've tried different versions and haven't seen any advantages.
Compared to illustrious it's better at prompt adherence with tag prompts, also it can understand natural language prompts since it has a proper llm as text encoder. Another thing it seems good at is producing high detailed images at medium resolution.
DiT is always better than Unet
@RisingV I've created a couple hundred images. I even tried the ones in the preview—and yeah, they're not quite right; they still require manual adjustments and resizing.
It might be useful for advertisers and anyone looking for an alternative to Banana, Flax, or Image.
Well nobody is forcing you to use it. And I am not saying it does not have it's downsides (for me that would be on the creators side for license reasons and finetuning capabilities). But I don't think it will replace larger models like the ones you mentioned. This is purely (or at least mainly) an anime model with huge character and style knowledge, so it's main competitors are sdxl based anime models like illustrious and noobai, also because they require similar computational ressources.
@Metisdark It is significantly better than illustrious, its not even close. The prompt adherence alone is way better. Lora training is faster and I think better at capturing identity than any Image 2 Image models. It only takes like 20 mins to train a lora and you get character identity solved, this is pretty huge.
@snowytechna857 So I'm here to figure out exactly what's good about it. I wrote that I’ve already experimented with it. I wrote that I’ve tried different projects, including from the preview. I still have to edit every piece of art as usual. And it still can’t handle complex compositions, multiple characters, and so on.
People tell me, “It’s awesome—if you don’t want to use it, don’t.” So show me some examples.
Could you consider adding Preview 3 to Civitai trainer?
I'm aware is just a preview not a final version and any lora trained with it will be a throw away once final release is out. But I don't think anyone minds that since you can always do another lora for the final version. I don't really want to deal with kohya in the meantime.
They need to buy a license, just like when they made it online here, they bought a license go to tensor art they have anima training
Great model! It actually made me believe in local generation again. Just one question: is it possible to use custom encoders, like Qwen 3, but with different parameter counts (e.g., 1.5B or 7B)?
I haven't tried it, so I can't say if and how good it works, but someone managed to make Qwen 3.5 4B work with it. Here is the model: https://civitai.com/models/2455272/anima-2b-qwen-35-4b-text-encoder
yes, but the quality is way worse.
BEST MODEL EVER!!!!!! WE NEED CONTROLNET!!!!!!!!!!!!!!!!
If you need ControlNet, take a look at the workflow I've uploaded.
@SAN0 ayyyyy thanks!
Details
Files
anima_preview3Base.safetensors
Mirrors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
animaOfficial_preview3Base.safetensors
anima-preview3-base.safetensors
animaOfficial_preview3Base.safetensors
anima-preview.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
animaOfficial_preview3Base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima_preview3Base.safetensors
anima_preview3Base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
anima-preview3-base.safetensors
Available On (3 platforms)
Same model published on other platforms. May have additional downloads or version variants.












