I. Introduction
AnimaYume is a text-to-image model fine-tuned from Anima, a high-quality anime-style image generation model developed by CircleStone Labs. It builds upon Cosmos 2, a model developed by NVIDIA’s research team.
II. Information
For version 0.1:
This model is a preview version fine-tuned from the Anima base model using a custom dataset. Training was conducted across multiple resolutions ranging from 768 to 1280 pixels, with a primary focus around 1024. The goal of this release is to improve stability and minimize unwanted artifacts when producing high-resolution images.
Notes: All the example images at this version were generated at the resolution 1024x1536 or 1536x1024
For version 0.2:
This model is a continuation of AnimeYume v0.1. In this version, I improved the quality of my dataset and used several techniques to prevent oversaturation and low-quality outputs. Based on my testing phase, I observed that the prompt coherence is better than v0.1, and the model remains very stable when generating images at a resolution of 1536.
Note: I am still waiting for the final version of Anima and testing some methods to make my training process faster. I know the license might make the model less popular, but I only care about whether the model is good or not. I’m aware that many others use better licenses, but I’m too lazy to spend a bunch of money training a model from scratch.
For version 0.25:
This version was trained on Anima Preview 2. Due to several issues with the base model, such as overfitting, black/white borders, quality inconsistencies, and problems with artist tags, I decided to focus primarily on improving the model’s knowledge, reducing these issues, and making it as stable as possible.
Note: In this version, I did not attempt to improve the model’s style. I tried doing so, but it caused the model to forget some of its existing knowledge. The training process is similar to v0.2, but the dataset has been adjusted to better address the issues present in Anima Preview 2.
For version 0.3:
This version was trained using Anima Preview 2. It is an experiment with a new training method for the model. You can consider it as another branch of AnimeYume 0.25, developed in parallel. However, this version uses new techniques and a larger dataset compared to v0.25.
Note: In this version, I experimented with a new training approach, so the model is slightly different from v0.25. Additionally, all example images were generated using prompts shared with users on CivitAI to evaluate whether this new method.
For version 0.4:
This version was trained on Anima Preview 3 using a custom dataset. In this release, I improved prompt understanding and artist style. Based on my testing, some artist styles match my expectations, although I haven’t tested everything in detail since I’m currently quite busy :<. Additionally, I fixed several issues from Anima Preview 3 that also appeared in Preview 2.
Note: I’ve only tested with simple test cases, not comprehensively, so if you encounter any issues, feel free to let me know. I also used a larger AI computing cluster to speed up the training process :D.
All example images were generated using prompts shared by users on CivitAI, as I wanted to evaluate the model’s performance.
For version 0.5:
This version was trained on Anima Base v1.0 using my custom dataset (a mix of a small e621 dataset and Danbooru). In this release, I added many new characters and improved the existing ones. I also enhanced support for various artist styles, allowing the model to generate results that are much closer to the original styles. In addition, the model now understands some concepts and knowledge from e621, although the support is still limited.
Notes: I’ve only tested the model with a few simple test cases so far, so if you encounter any issues, feel free to let me know. This release can be considered a demo version showcasing my new training method, which focuses on preserving existing knowledge while adding new knowledge at the same time. The release also came sooner because I was finally able to use all the resources I had available :D
All example images were generated using prompts shared by users on CivitAI, as I wanted to evaluate the model’s performance using real user prompts.
For version 1.0:
This version was fine-tuned on Anima Base v1.0 using a variety of datasets, including Danbooru, e621, Gelbooru, and Konachan. The training approach differs from the original model in several ways. Most notably, I did not include quality score tags in the training data.
I also experimented with multiple captioning styles, ranging from traditional tag-based annotations to different forms of natural language descriptions, similar to the approach I used for Netayume Lumina.
Note:
This release has two versions, each trained using different methods. For the v1.0 Demo, I experimented with a mixed training approach, but it was difficult to control. The v1.0 Final is different from the v1.0 Demo, so please don't compare them directly, even though they were trained on the same dataset. I created both versions to test different ideas for training diffusion models. Since Anima is small enough, it gives me the flexibility to experiment with various training methods and see what works best.
Unlike Netayume, I did not use my full dataset for training. (The complete dataset contains around 25 million images, and if I were to use the entire dataset, I would train a model from scratch rather than fine-tune an existing one.) Additionally, I have been quite busy recently, so I have not been able to test this model as extensively as I would like. Moreover, this model dont have any default style :L
V1.0 final has a watermark in the model, which is not effect the results of generated images. Moreover v1.0 currently support chain of though prompt as you can see in my example images i did it
If you encounter any issues or have feedback, please feel free to share them with me.
All example images were generated using prompts provided by CivitAI users, as I wanted to evaluate the model's performance under real-world prompting conditions.
III. File Information
This file contains only the diffusion model and does not include a VAE or text encoder. To use it properly, you will need to download those components from the link here
IV. Notes & Feedback
This is an experimental fine-tuned release, and I am waiting for the final version release to tune it :D
Your feedback, suggestions, and creative prompt ideas are always welcome, every contribution helps make this model even better!
V. Acknowledgments
Big thanks to narugo1992 for the dataset contributions.
Credit to Circlestone Labs and Nvidia for the fantastic base model architecture.
If you'd like to support my work, you can do so through Ko-fi!
Description
FAQ
Comments (63)
Hi everyone, I’ve released a new version. In this release, I didn’t have enough time to test the model in detail because I’m currently very busy. Based on my testing, when comparing my model with Anima Preview 3, I feel my model might perform better (this is based on my personal preference).
Here is the comparison post, each image includes the prompt I used for testing:
https://civitai.com/posts/27912586
If you don’t mind, please give it a try and share your feedback. Thanks!
From my testing, when I test POV (especially girl POV), the success rate is quite low around 20% .
The prompt adherence also seems to have decreased a bit. Is it just me?
I was able to get good results using the base Anima P3, but yeah, higher resolution looks a lot cleaner,need more testing u guess
@Seii1 Hi, Would you mind telling more details i tested about 50 images generated with tag pov and didnt see any problem?
Is this preview 3? I assume not, still will give it a ago :D
@GPUPoorChad Yes 0.4 was trained on preview 3
Hi, may I ask what trainer you used for the fine-tuning? Thank you!
@qizongzui The model was trained using a custom script, derived from the original sd-scripts repository
amazing model but you need to fix pixelated lines
Oh i generated about 300 images and evaluated by my self and gemini and didnt see the problem 😅😅. Would you mind telling more details
@duongve13112002 NVM, i forgot to add new version to comfy XD, your model is so great get all my points
You can improve the quality by changing the VAE to a different one. I recommend using the QwenimageVAE-Liquid v7 , you can find it here: https://civitai.com/models/2487530/qwenimagevaeliquid1087
@danque what about samplers and scheduler?
@Neon_signs What about them? I am using mostly ResMultistep (sampler) with normal (schedule), or Res_2m with bong_tangent (this one is not a default)
@danque yeah that's the thing I wanted to know lol,much appreciated
Great work. Compared to version 0.3, aliasing(also known as dithering) is noticeably reduced.
People sometimes complain 'without any image or prompt information' and it's good for the mind not to care.
V0.4 is just so awesome,great work, only problem I noticed so far is that it really fights to create dim lighting,dark scenarios somehow
Thank you for another great checkpoint.
Tested with character Lora - V0.4 basically takes Preview3 outputs and just makes it better (posted comparison in gallery).
Your model is great! I hope that in the future Civitai will agree to a commercial license with Circlestone-Labs for online generation of fine-tuned checkpoints and training of Lora models. But for now, locally, your model produces incredible results. 💖
Mostly amazing, but suddenly became way worse with wrong number of fingers and toes, also as others say - problems with dim lighting/overbrightened subject.
I mean, it correctly draws dim light and heavy shadows on the background but leaves characters brightly lit.
Oh, let me find the root cause. This may be related to the style-tuned phases. Because the pre-trained base i tested i didnt meet these problems
@duongve13112002 It sometimes works fine but other times simply refuses to not overbrighten, I wish I could tell more, but I fail to pinpoint why it does that. Also, the part about fingers and toes: it is not extremely worse, but noticeably worse, and especially in unusual poses. Got absolute worst cases with barefoot characters kneeling, with a view from behind. Earlier versions/base Anima seem to be doing better in that one regard. But nevertheless, amazing model, thank you for your work!
Hey, I’m not completely sure, but I’ve observed that WAI-ANIMA generates images that are about 80% similar to this model. The style and background are slightly different, but not by much. I used the same settings for both models, so I’m not sure why this is happening.
All of the fine tunes are underbaked, they are new. They will diverge and become more different from the base model as people continue training. The differences on many of them are subtle atm.
@Drakeni No, I don’t believe that’s the reason. I ran the same test using AnimaYume v0.4 and Anima Preview 3, and the differences can be clearly observed.
@thaimannguyen4672 Honestly, I tested WAI Anima and found it to be a solid merged model. As for the similarities, I think that’s expected merged models often carries over shared traits from base model. From my perspective, this approach is a valid way to refine the base model, as long as it improves overall stability.
and ignore original model license, hide all credits, pretend it's a completely different model
Wai v0.1 was a mix of 0.5 animayume v0.4 + 0.5 anima p2.
@reakaakasky Let me add some context for those might wonder and reacted with laugh emoji.
This guy don't train model, all his models are merges. He don't give credits to original models, who spend money and time to train. He pretends it's his own model, with more training data, steps.
Wai sdxl's model is Noobai, but he deliberately conceals the truth, violating Noobai's license. After being questioned, he lied and said his model was based on Illustrious v1, which is bs, because noobai team calculated matrix similarity and triangulated his models between NoobAI v0.75 and v1. It's a merge of NoobAI, there is no Illustrious v1 in the model.
The biggest thief and liar is back.
@ikekph5 any wai sdxl alternatives do you highly recommended?
I wonder, Would you like a version with native 1536×1536 resolution support? Give me your idea :D
i think qwen vae is limited, flux2 vae is better, and klein is very handy even do 2k upscale + denoise, it can directly denoise from t 0.7 to 0 in one step
@reakaakasky I think i will test it, Anyway this is just an experiement stuff :D
Anima itself is still not there yet is it? The preview 3 made improvements but became harder to use. Might take a few months still. Great model though. It would be amazing if it could, the VAE is so much better than the 4 channel sdxl ones...
@BoundingBoxes what do you mean harder at use? For me Animeyume v0.1 and v0.4 the same simple using...
Учитывая, что текущая модель отлично из коробки поддерживает 1024*1536, как будто не будет проблемы масштабироваться до 1536*1536
Анима выглядит перспективно и интересно
Just a question - could be this one be GGUFed with MTP support to get more s/it ?
@reakaakasky Isn't the flux vae censored to hell and back?
If it possible, can you share script for model trainings and/or lora tanning
Actually, it depends on the dataset, so there’s no fixed configuration for it. For LoRA, I recommend using rank 16 and alpha 8 for character training and rank 32 with style, with a learning rate of 1e-4 using the AdamW optimizer, and keeping everything else as the default settings in the Kohya repo.
It's still the best fine-tuned model. It combines the stability and the diverse output quality of the official version, and its natural language processing capability remains strong!
Seeing this person could make a decent equivalent to NoobAI-XL (NAI-XL) (A model with the ability of responding to prompts made with tags used in Danbooru and/or in e621) based in Neta Lumina, I find weird duongve13112002 haven't tried do the same with Anima yet.
Hi, currently I don’t want to spend too much compute on the preview version (since ComfyUI sponsors this model, training it now would be a awful idea and i have other projects which need high compute right now). My model was trained on a small subset of my dataset (my current dataset is updated to 25/4/2026). At the moment, I’m also focusing on preserving the base model’s knowledge while improving some areas the model has forgotten.
https://civitai.red/models/2544636/wai-anima?modelVersionId=2859702
This model integrates your base model without any mention of it on the release page, and I think this is inappropriate.
Metadata can be read in ComfyUI, and the merged parts can be restored through differences.
hmm, seems this time he forgot to remove the metadata.
Unfortunately,
99% of users, even if they know this, won't care.
98% of users won't understand what you're talking about.
97% of users believe he is the biggest trainer.
And if you mention this in their server, you'll be banned immediately.
Reality is harsh.
what can we do to punish stealer
@anxiousxeon Unfortunately, nothing. Most people just indulge in that kind of aesthetic and reject those who question WAI. All I can do is inform
For those who wondering, yes, you've found one of the real Gigachad trainer, who added anima training support in kohya-ss/sd-scripts, has full danbooru dataset, and probably private GPU cloud with serval clusters.
Please support the real model trainers who have spent time and computing resources, instead of those "model thieves" who only steal and merge other people's models, deliberately hide credit from the original authors, pretend those are their own models.
You can easily distinguish between a real trainer and a thief. A real trainer shares basic info about dataset, training parameters, even own training script. Actives on dev platforms, hf, github etc.
A thief only has monetization accounts on AI platforms. You know nothing about their models because they can't let you know.
Thank you for all the hard work you and other great creators like @reakaakasky have done on all the preview versions of this great anime model anima. The fine-tunes and stability really have made the model a joy to use. Thank you 😊 🙏
@duongve13112002
Бро, там Анима-в1-База вышла! Жду новую версию твоей модели!
WE ARE WAITING FOR BASED ANIMAYUME
Be patient please it's coming soon 🙏
Hi, the training process has started, and it will take longer than I expected because I want to experiment with some things (including adding some e621 datasets :L - it might break the model, but let’s try anyway =)) ). So, the next version based on Base 1.0 will take longer than expected :D
Thanks, king. I'm using only your anima model, it works so good
Thank you for the update. Can't wait to try to use it with base ❤️
While we wait for Base release I was wondering if you helped implement Anima training in sd-scripts? I saw your name on the changelog.
As a sd-scripts enjoyer I wanted to give you a massive thank you for the effort.
He did the port from diffusion-pipe, yes. I'll take advantage on this comment to report that there's some strange memory allocation behaviour with this port.
1. On windows higher batch sizes than 1 allocate memory to the shared memory on RAM making the trainings very slow.
2. On linux during the training I run out of memory after an arbitrary number of epochs with higher batch sizes, there seems to be some kind of spike that's not there on sdxl trainings.
3. Once I had a training that I had successfully ran with batch size 10 no longer be able to do that.
I don't' have more details than this, sadly.
@IndolentCat hi would you mind telling more detail about your pc which you used for training, i tested with gpu with 16 GB VRAM it worked normally on training lora.
@duongve13112002 I can't help with that on the Windows one. It happened to a few people at Arc en Ciel discord. I specifically remember KojimBomber/Nullptr showing screenshots. Edit: He says diffusion pipe does the same, so disregard.
The batch size problem happened on Linux at runpod. Yeah, variable environment and all that, but the fact that it never ever happened with SDXL still makes it suspicious. So does the way that it happens after a random epoch, it could be the 2/10, it could be the 6/10, etc.
Not sure about windows, but on linux try to add this to your startup script
export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
Also, I don't think Cosmos need higher batch size > 1. The GPU is already saturated with bs 1 (unless you are using h100 or other beast). Try "gradient_accumulation".
does anyone have a full list of the trained characters for this
Looking forward to the latest version of Anima from the guru.
期待大佬最新微调的Base模型



















