I. Introduction
AnimaYume is a text-to-image model fine-tuned from Anima, a high-quality anime-style image generation model developed by CircleStone Labs. It builds upon Cosmos 2, a model developed by NVIDIA’s research team.
II. Information
For version 0.1:
This model is a preview version fine-tuned from the Anima base model using a custom dataset. Training was conducted across multiple resolutions ranging from 768 to 1280 pixels, with a primary focus around 1024. The goal of this release is to improve stability and minimize unwanted artifacts when producing high-resolution images.
Notes: All the example images at this version were generated at the resolution 1024x1536 or 1536x1024
For version 0.2:
This model is a continuation of AnimeYume v0.1. In this version, I improved the quality of my dataset and used several techniques to prevent oversaturation and low-quality outputs. Based on my testing phase, I observed that the prompt coherence is better than v0.1, and the model remains very stable when generating images at a resolution of 1536.
Note: I am still waiting for the final version of Anima and testing some methods to make my training process faster. I know the license might make the model less popular, but I only care about whether the model is good or not. I’m aware that many others use better licenses, but I’m too lazy to spend a bunch of money training a model from scratch.
For version 0.25:
This version was trained on Anima Preview 2. Due to several issues with the base model, such as overfitting, black/white borders, quality inconsistencies, and problems with artist tags, I decided to focus primarily on improving the model’s knowledge, reducing these issues, and making it as stable as possible.
Note: In this version, I did not attempt to improve the model’s style. I tried doing so, but it caused the model to forget some of its existing knowledge. The training process is similar to v0.2, but the dataset has been adjusted to better address the issues present in Anima Preview 2.
For version 0.3:
This version was trained using Anima Preview 2. It is an experiment with a new training method for the model. You can consider it as another branch of AnimeYume 0.25, developed in parallel. However, this version uses new techniques and a larger dataset compared to v0.25.
Note: In this version, I experimented with a new training approach, so the model is slightly different from v0.25. Additionally, all example images were generated using prompts shared with users on CivitAI to evaluate whether this new method.
For version 0.4:
This version was trained on Anima Preview 3 using a custom dataset. In this release, I improved prompt understanding and artist style. Based on my testing, some artist styles match my expectations, although I haven’t tested everything in detail since I’m currently quite busy :<. Additionally, I fixed several issues from Anima Preview 3 that also appeared in Preview 2.
Note: I’ve only tested with simple test cases, not comprehensively, so if you encounter any issues, feel free to let me know. I also used a larger AI computing cluster to speed up the training process :D.
All example images were generated using prompts shared by users on CivitAI, as I wanted to evaluate the model’s performance.
For version 0.5:
This version was trained on Anima Base v1.0 using my custom dataset (a mix of a small e621 dataset and Danbooru). In this release, I added many new characters and improved the existing ones. I also enhanced support for various artist styles, allowing the model to generate results that are much closer to the original styles. In addition, the model now understands some concepts and knowledge from e621, although the support is still limited.
Notes: I’ve only tested the model with a few simple test cases so far, so if you encounter any issues, feel free to let me know. This release can be considered a demo version showcasing my new training method, which focuses on preserving existing knowledge while adding new knowledge at the same time. The release also came sooner because I was finally able to use all the resources I had available :D
All example images were generated using prompts shared by users on CivitAI, as I wanted to evaluate the model’s performance using real user prompts.
For version 1.0:
This version was fine-tuned on Anima Base v1.0 using a variety of datasets, including Danbooru, e621, Gelbooru, and Konachan. The training approach differs from the original model in several ways. Most notably, I did not include quality score tags in the training data.
I also experimented with multiple captioning styles, ranging from traditional tag-based annotations to different forms of natural language descriptions, similar to the approach I used for Netayume Lumina.
Note:
This release has two versions, each trained using different methods. For the v1.0 Demo, I experimented with a mixed training approach, but it was difficult to control. The v1.0 Final is different from the v1.0 Demo, so please don't compare them directly, even though they were trained on the same dataset. I created both versions to test different ideas for training diffusion models. Since Anima is small enough, it gives me the flexibility to experiment with various training methods and see what works best.
Unlike Netayume, I did not use my full dataset for training. (The complete dataset contains around 25 million images, and if I were to use the entire dataset, I would train a model from scratch rather than fine-tune an existing one.) Additionally, I have been quite busy recently, so I have not been able to test this model as extensively as I would like. Moreover, this model dont have any default style :L
V1.0 final has a watermark in the model, which is not effect the results of generated images. Moreover v1.0 currently support chain of though prompt as you can see in my example images i did it
If you encounter any issues or have feedback, please feel free to share them with me.
All example images were generated using prompts provided by CivitAI users, as I wanted to evaluate the model's performance under real-world prompting conditions.
III. File Information
This file contains only the diffusion model and does not include a VAE or text encoder. To use it properly, you will need to download those components from the link here
IV. Notes & Feedback
This is an experimental fine-tuned release, and I am waiting for the final version release to tune it :D
Your feedback, suggestions, and creative prompt ideas are always welcome, every contribution helps make this model even better!
V. Acknowledgments
Big thanks to narugo1992 for the dataset contributions.
Credit to Circlestone Labs and Nvidia for the fantastic base model architecture.
If you'd like to support my work, you can do so through Ko-fi!
Description
FAQ
Comments (44)
It's here!
and the cameo prompts from other finetunes + Lizardon1205 haha
What kind of e621 knowledge is included?
About some artists and tags. I just random choice from mydataset, i will check it later :<
Yay will test it later
I love wataten!!
Finally! Animayume v0.4 was amazing, Can't wait to see what the new version is capable of. Also is there a way to know all the new characters added?
From v1 and preview 1 to the base and v5. Thank you for all your hard work and great Checkpoints and LORAs. Your a legend.
This one and dasiwa are currently the best finetune anima models
dasiwa models is poor cuz It looks very similar to Pornmaster's models featuring chubby characters aged 30 and up,
Animayume models much better and follow anime prompts whithout any output changes
@ollymolly20073840 dasiwa perform better with multiple characters. Yume is better with a couple or less
0.5 is excellent. Even works well with LoRAs trained on 0.4.
I hope to see the new characters up to today.
Awesome, been waiting for v0.5!, btw is there a possibility to know Wich new characters did you add?
I’d appreciate it if you could list the new characters, if it’s not too much trouble.
Hi, I want to clarify the dataset. I updated the Danbooru dataset to May 1, 2026, and the e621 dataset to April 25, 2026 (for e621, I used a very small dataset).
As for the new characters, many of them come from recent anime and other internet sources (for example, CivChan). Basically, all of them already have tags on Danbooru, so you can check them there. Because there are too much :((
This version is an experiment when i using a small subset of my dataset. The final may be longer because i can not use my resource at right now :<
Do we need any special tagging for the new characters? Also are all the characters included up to May of 2026 or do they need to have an specific number of tags for them to work?
I've tried a few "newer" characters including Velina Airgid from zenless zone zero and Label from nikke and the results are pretty inconsistent even with heavy tag support. So it seems they might not be fully working or I'm missing something.
@GoonetteAI_ Oh i think the character need at least 300 images may be work correctly. Moreover, i am not use full danbooru because i have to select the image has the best quality + filter duplicate too
Thank you for the great model.
If possible, could you consider expanding Blue Archive coverage in future dataset updates? It remains one of the most popular franchises on Booru sites. For instance, tags such as kei_(student)_(blue_archive) and rio_(armed)_(blue_archive) have existed since last year but are currently missing from both the Anima base and AnimaYume datasets.
Appreciate your efforts!
Hey, Recently I have seen many people switching to Anima. Is Anima better than illustrious?
@w4ifuzdeity751 It is better. Anything about SDXL is just too old.
Your model works wonderfully with my LoRAs, thank you! 🙌✨
Thank you for another amazing checkpoint.
However, from the testing I did with Base trained character Loras, using v0.5 for inference gives me less of a quality bump compared to previous versions.
I tried going back to v0.4 and its giving me better results [more detailed, better aesthetic] than v0.5.
Not sure if this is expected or if its just me. Will continue testing though.
Hi this is because v0.4 was trained in two phases ( include aesthetic with reinforcement learning) but v0.5 only with reinforcement learning with large dataset compare to v0.4, so it can know many new character and help training lora more better
@duongve13112002 So, future release will be trained in 2 phases, right?
having fun with the model thank you
If you don’t mind me asking, what training parameter settings did you use?
Yes of course it is simple i use sd-script implement with my custom modification. For the training i choose learning rate is 5e-5 (i trained llm adapter with lr 5e-7, i think uou shouldnt train adapter if you dont have large dataset) with adamw, scheduler is cosine with restart. Shift = 1 and sampling is sigmoid.
I hope this solves the Anima issue of images using the older dataset of artists. I thought it was just a illustrious issue because those models are old
Seems to want to want to do ass focus constantly when I don't want it to, and I just want character lying on stomach facing the camera, so there is clearly some bias in your training set compared to base. Overall though, good model!
It knows you are an "ass" man xD
Hi i wonder, do you want a huge improvement. 🤔🤔🤔 (my dataset currently now 25M bruh). This can be a large training and can take a huge money bruh. I am still thinking because my budget is limit right now....
finetune PixelDiT
@reakaakasky ..... More resource 🤣🤣🤣
Of course we want!!!
@duongve13112002 or finetune an upcoming Cosmos 3 Edge 4B 😜
@Lynx2025 or ideogram4 v:
@duongve13112002 nice work!!, artist Jabcomix is trained in model animayume?
Are you planning on making an artist list like you did with netayume lumina? I'd like to know what artists the model know
I'm getting both worse visual quality and worse prompt following in 0.5, compared to 0.4 🤔, gonna keep using the old one.
Hi everyone,
Please pay attention to the current Anima checkpoints. Based on my testing, most of them are extremely similar to the original Anima base model, with only very minor changes.
You can easily verify this by comparing the model weights layer by layer against the base model. The cosine similarity is nearly 0.9999, while the L1 distance difference is very small. In other words, the checkpoints are almost identical to the base model from a weight perspective.
I also compared them using the exact same prompts (both simple and complex) under identical settings. The generated images are nearly indistinguishable from each other. The behavior is very similar to using a LoRA on top of the base model rather than a significantly different checkpoint.
As for the models I've recommended, I've personally tested them and found that they deliver noticeably better results.
I'm not trying to start a conflict with anyone. My goal is simply to provide transparent information and help the community better understand the current state of these models. In my opinion, releasing checkpoints with such minimal differences only adds confusion and contributes to the growing amount of low-quality model variants in the ecosystem.
😑😑😑😑😑
I tried to get chatgpt to write a comfy node that could compare model similarity. And I failed; I don't know how to guide chatgpt, because I don't know what exactly I need. Skill issues 😭
Is there any chance you can make it to a comfy node or something easy to use? Much appreciated. And it would be really helpful.
That's the end of game of anything with AI. Once everyone can do something you have so much slops that nobody can tell what to look for, killing itself.
I have to say that after tuning the weight of negative prompts, the Yume model delivers outstanding visual results. It retains intricate color and shadow details even in complex multi-character compositions. I’ve tested numerous other models, most of which tone down this richness—their outputs end up looking quite similar to ill (though anima’s level of detail is far superior by a wide margin).
That said, this fine-tuning works against goals of heightened fine detail or authentic hand-drawn aesthetics, leaving images with an obvious artificial AI look.
On the downside, Yume is far more prone to distorted hands and feet in busy scenes. This might be a trade-off stemming from its strong ability to render complex shadow hues. I hope future iterations can fix limb deformities while preserving its current artistic style.
Additionally, the underlying Anima base dataset seems to be the culprit behind another flaw: many tagged artist styles correspond to very old-school art aesthetics, which often produce subpar, unsatisfactory renders.Therefore, it is currently recommended to use Yume in combination with style LoRAs.
All things considered, this is an excellent piece of work, and I’m eagerly looking forward to the updated releases!
lovely



















