CivArchive
    Preview 139395821

    🎨 See-Krea2-Genesis V2

    More detail. Richer scenes. A personal pixel-art look.

    Krea 2 Turbo · 8-step workflow · Natural-language prompting · Qwen3-VL
    Anime illustration · Detailed characters and backgrounds · Pixel art · ComfyUI


    ✨ What is See-Krea2-Genesis?

    See-Krea2-Genesis is my personal anime checkpoint built on Krea 2 Turbo. The aim is to make anime illustration the model’s natural visual direction while keeping the fast, straightforward Turbo workflow.

    V2 builds on the original Genesis with longer training at higher resolution. The focus is on more image detail, better prompt following, detailed anime women and richer backgrounds. I also trained in my own pixel look, two character triggers, and a retro handheld-game direction using my own Game Boy captures.


    🆕 What changed in V2?

    • Longer training at higher resolution. V2 received more training with higher-resolution material than V1.

    • More detail in the image. The training included detailed anime women and detailed backgrounds, with attention to both the character and the surrounding scene.

    • Better prompt following. V2 follows the scene and character details in your description more closely.

    • Two trained triggers: seesee_elf and seesee_kitsune.

    • My own pixel-art look. It can be requested with different phrases, including a simple pixel art cue.

    • Game Boy-inspired imagery. My own Game Boy captures were included in the training data.

    📝 About English, Japanese and manga text:
    The training included English and Japanese text as well as manga material. I have not confirmed an improvement in readable text generation. Please treat spelling, lettering and longer text as something to test, rather than a promised V2 feature.


    🧝 Trained character triggers

    These two trigger words were included during training. Use the exact spelling in your positive prompt, then describe the scene, pose, clothing and lighting you want.

    🧝 seesee_elf

    Trigger: seesee_elf

    EXAMPLE 01 · SEESEE_ELF

    Example prompt

    eesee_elf, A single female character stands leaning against a glowing pink vending machine in a neon-lit urban alleyway. The image is rendered in a high-quality, semi-realistic digital anime style with sharp focus and vibrant color grading. The framing is a full-body vertical shot, capturing the subject from head to toe with a slight low-angle perspective. She leans her left hip against the machine, her left hand tucked into her skirt pocket, while her right hand holds a pink beverage can at chest level. Her gaze is directed forward, looking slightly past the camera. She has long, straight platinum blonde hair that falls past her waist, pointed elf ears, and large amber eyes. Her face is pale with a neutral, calm expression. Her body is slender with long legs covered in sheer black pantyhose. She wears a cropped pink blazer over a black bandeau top, a matching high-waisted pink mini skirt, and black platform high-heeled shoes with ankle straps. A black lace choker and a silver necklace with a cross pendant adorn her neck, and long black earrings dangle from her ears. The background features a wet, reflective cobblestone street and blurred neon signs, including a vertical pink sign reading "OPEN" and a purple sign reading "BAR". The vending machine displays Japanese text "ドリンク" at the top and contains rows of colorful bottles. The lighting is dominated by the intense pink glow from the vending machine, which casts a strong magenta hue over the scene and creates sharp reflections on the wet ground.

    🦊 seesee_kitsune

    Trigger: seesee_kitsune

    EXAMPLE 02 · SEESEE_KITSUNE

    Example prompt

    seesee_kitsune, A single female character with fox ears and a large, fluffy tail sits on an ornate wooden chair, her legs spread wide apart. The scene is rendered in a high-fidelity anime style, blending digital illustration with realistic textures and lighting. The framing is a full-body shot, capturing the character from head to toe, centered within the composition. She leans forward slightly, resting her arms on the back of the chair, looking directly at the viewer with a confident, subtle smile. Her face features sharp, angular features, bright orange eyes, and long, flowing orange hair that cascades down her back. Her body is slender, with visible curves and a defined waist. She wears a black, form-fitting outfit adorned with intricate gold chains and a prominent red gemstone at the center of her chest. Her legs are clad in sheer black stockings, and she wears black high-heeled boots with gold accents. The background is a dimly lit room with a vintage aesthetic, featuring a large mirror with light bulbs, a table with candles, and a vase of red roses. The lighting is warm and dramatic, casting soft shadows and highlighting the character’s features and the rich textures of her clothing and surroundings.


    👾 My pixel-art look

    V2 includes my own pixel look. You can ask for it in different ways; the prompt should contain pixel or 16Bit. A simple pixel art cue is a useful starting point. Combine the style cue with a normal description of the character or scene.

    Style cues to try: pixel art · a detailed pixel illustration · 16Bit pixel art

    Pixel look · Example 1

    EXAMPLE 03 · PIXEL-ART

    Example prompt

    A solitary figure sits in a command chair facing a vast spaceship viewport, rendered in pixel art with a visible grid of square pixels, a limited palette of neon blues, magentas, and reds, and hard-edged blocky shapes. The camera angle is a symmetrical, eye-level rear view from behind the pilot, looking forward into the cockpit. The figure is a dark silhouette with a rounded head, seated centrally and facing away from the viewer toward the controls. The cockpit interior is dense with angular consoles on the left and right, featuring small glowing blue screens and clusters of red, orange, and green square buttons. The large, multi-paned window reveals outer space, dominated by a large blue planet with bright rings on the left and a swirling nebula of pink, purple, and orange on the right. A small, cyan wireframe sphere floats in the nebula. The floor is dark blue with red rectangular lights, and thick, segmented pipes run along the bottom corners of the frame.

    Pixel look · Example 2

    EXAMPLE 04 · rendered in pixel art with a visible grid of square pixels

    Example prompt

    A single anime-style girl kneels in a dark, narrow urban alleyway, rendered in pixel art with a visible grid of square pixels, a limited color palette, and hard-edged blocky shapes. The camera is positioned at a low angle, looking up at the figure who is facing forward with her body angled slightly to the left. She has short, light blue hair with white highlights, large expressive eyes with pink and yellow irises, and prominent white cat ears on top of her head. She wears a black long-sleeved top, a short black pleated skirt, and black thigh-high stockings. Her legs are bent at the knees, with her thighs spread apart, and she wears brown shoes. The background features dark building walls, a glowing red lantern on the left, and scattered bokeh lights in the distance, all defined by the chunky, low-resolution aesthetic of the medium.


    🎮 Game Boy-inspired scenes

    I also included my own Game Boy captures in the dataset. To ask for a retro handheld-game scene, try a description like this:

    a retro 8-bit video game level, rendered in Game Boy Color era pixel art with a visible coarse pixel grid

    Use that as a style direction, then describe the level, setting and objects you want to see.

    EXAMPLE 05 · GAME BOY-STYLE

    Example prompt

    [A blue cat character stands on a grassy ledge in a retro 8-bit video game level, rendered in Game Boy Color era pixel art with a visible coarse pixel grid, small resolution, and limited color palette. The framing is a horizontal side-scrolling view showing a light blue sky with white clouds, a distant green mountain range, and a tan sandy ground area. The blue cat faces forward with a simple closed-eye expression, positioned near small orange flowers and green bushes on the right. Floating above are two yellow, round creatures with green accents, one resting on a small floating platform and another hovering in the upper right corner. A brown wooden crate sits on the far left edge, and a blue water hazard spans the bottom left. The bottom of the screen displays a black HUD bar with white pixelated text reading "WORLD 1", a counter showing "x4", and a timer reading "T:79".


    📦 Which version should I download?

    ⭐ Genesis-int8_convrot

    Size: 13.2 GB
    Weight error: 0.99%

    Recommended for: Best overall quality.

    Works on any GPU that supports INT8, including RTX 30-series and newer.

    ➡️ Choose this version if you are unsure which one to download.


    Genesis-fp8_scaled

    Size: 13.1 GB
    Weight error: 2.66%

    Recommended for: RTX 40-series (Ada) and newer GPUs.

    Uses the native FP8 path and is a good choice for GPUs with strong FP8 support.


    Genesis-int8

    Size: 13.2 GB
    Weight error: 1.55%

    Recommended for: Users who want INT8 with slightly faster runtime performance.

    Similar size and quality to Genesis-int8_convrot, but optimized more toward inference speed.


    Genesis-nvfp4

    Size: 7.7 GB
    Weight error: 9.41%

    Recommended for: RTX 50-series / Blackwell GPUs.

    Uses native NVFP4 and requires significantly less storage and VRAM than the INT8/FP8 versions.

    ➡️ A great option if you have a Blackwell GPU and want a much smaller model.


    Genesis-int4_convrot

    Size: 6.9 GB
    Weight error: 16.43%

    Recommended for: Low-VRAM systems or users who want the smallest possible download.

    This is the smallest Genesis build, but it also has the highest quantization error.

    ➡️ Choose this version when VRAM or disk space matters more than maximum quality.


    ⚠️ How to read the "Weight Error"

    Lower is better.

    The weight error indicates how much the quantized model weights differ from the original model.

    • 0.99% → Very close to the original weights

    • 1–3% → Excellent quality / very small difference

    • ~9% → More aggressive compression

    • ~16% → Significant compression, mainly intended to save VRAM and disk space

    A higher weight error does not automatically mean that generated images will be worse by the same percentage. It only describes the difference between the original and quantized model weights.

    Quick recommendation

    ⭐ Best quality: Genesis-int8_convrot
    ⚡ RTX 40-series: Genesis-fp8_scaled
    🚀 Faster INT8: Genesis-int8
    💾 RTX 50-series / smaller model: Genesis-nvfp4
    🪶 Smallest possible version: Genesis-int4_convrot

    ⚠️ Read the "weight error" column correctly

    That number is the measured deviation of the weights from the BF16 master, not of the images. The two are only loosely related.

    In side-by-side testing all five builds look good, and the differences are small enough that you generally have to put two images next to each other to spot them. So do not rule out the small builds because of the number — if 6.9 GB is what fits your card, try it before assuming it is a downgrade.

    What the numbers are actually good for is ranking the formats against each other, and explaining why int8_convrot beats fp8_scaled at the same file size: FP8 E4M3 has three mantissa bits, INT8 has 256 evenly spaced steps, and the ConvRot rotation spreads outliers before quantising instead of letting one large value eat the scale.



    The settings I personally use and recommend as a starting point:

    Steps:       8
    CFG:         1.0
    Sampler:     euler
    Scheduler:   simple
    Denoise:     1.0
    Resolution:  ~1 MP, for example 1024 × 1024

    Sampler alternatives

    • euler + simple — my default. Neutral, clean, predictable.

    • er_sde + simple — slightly more texture and character, works well at 10–12 steps.

    ⚠️ CFG must stay at 1.0

    This is a turbo model. Raising CFG above 1 will burn the image — that is not a style choice, it breaks.

    Two consequences people trip over:

    • Your negative prompt does nothing. At CFG 1.0 the unconditional branch is never computed, so negative text has no mathematical effect on the result. Whatever you put there is ignored.

    • Exclusions must be phrased positively. If you do not want gloves, write bare hands in the positive prompt. Writing gloves in the negative field changes nothing.


    💡 Prompting guide

    Natural-language descriptions are my preferred way to prompt this model. Start with what the image should show, then add the details that matter most.

    A useful order:
    Medium and style → camera and framing → pose and action → face and hair → clothing → background → lighting.

    • Write clear sentences. You do not need to hit a particular word count; add detail when it gives the model useful information.

    • Name the framing. For example: close-up portrait, low-angle shot or wide full-body composition.

    • Be specific about hands and actions. Describe each hand separately when the pose matters.

    • Describe materials and clothing layers. This gives the model more to work with than a clothing label alone.

    • Give the background its own description. V2’s training included detailed environments as well as detailed characters.

    • Keep the subject clear. Consistent pronouns and an explicit character count help avoid ambiguity.

    • For the pixel look, include pixel or 16Bit. For the trained character cues, use seesee_elf or seesee_kitsune exactly.

    If a detail is missing, first make the wording more specific instead of relying on long tag lists or weighting syntax.


    🖼️ Example Prompts

    Five prompts covering very different territory. Each one is written the way the model likes to be talked to — use them as templates and swap the content.

    1 · Dark sci-fi portrait

    A striking dark sci-fi anime portrait shows a cybernetic girl against a fractured
    digital void. The illustration uses high-contrast monochrome tones with glowing cyan
    glitch accents, precise line art and a cold surreal aesthetic. A tight close-up frames
    her head and shoulders from a slightly low angle. She holds perfectly still, chin
    lifted, staring past the viewer. Her short white hair is cut in a sharp asymmetric bob
    with one side shaved, and thin luminous seams trace her cheekbone and temple. Her eyes
    are solid cyan with no visible pupil. She wears a matte black high-collared bodysuit
    with exposed cabling along the neck. Fragments of broken screen geometry drift behind
    her. Rim lighting picks out the edge of her jaw against the darkness.

    2 · Neon graphic full-body

    A stylised graphic anime artwork depicts a mischievous spirit girl rendered entirely
    in liquid chrome, outlined in electric magenta. The full-body composition places her
    mid-leap above a glowing hexagonal sigil, one arm thrown back and the other reaching
    forward with fingers splayed, grinning with her eyes squeezed shut. Her form is a
    glossy silver silhouette with flowing hair that breaks apart into floating droplets.
    High-contrast flat colour, thick neon outlines and hard shadow shapes define the style.
    The background is deep violet with radiating light streaks. Magenta reflections pool
    beneath her on a dark mirrored floor.

    3 · Fantasy scene with heavy background

    An anime illustration shows a sorceress standing at the centre of a flooded jungle
    temple, framed by enormous carved stone faces half-swallowed by roots. A wide full-body
    shot from slightly below places her small against the architecture. She stands ankle
    deep in still water, one arm outstretched with palm turned upward, looking calmly
    toward the viewer. Her long moss-green hair falls past her waist and her eyes glow
    pale gold. She wears layered ivory robes bound with braided cord, a heavy bronze
    pectoral collar and bare feet. Vines and hanging orchids drape the ruins behind her.
    Shafts of warm daylight break through the canopy and scatter across the water surface.

    4 · Slice of life

    A soft anime illustration of a young florist arranging a bouquet on a wooden workbench
    in a small sunlit shop. A relaxed medium shot from across the counter catches her
    mid-motion, holding a stem of white ranunculus in one hand and secateurs in the other,
    head tilted in concentration. Her wavy honey-blonde hair is pinned up loosely with
    a few strands escaping, and she has warm grey eyes and a light scattering of freckles.
    She wears a soft blue linen shirt with the sleeves rolled to the elbow under a canvas
    apron marked with plant stains. Buckets of flowers, brown paper and twine crowd the
    bench around her. Late afternoon light comes through the shop window and warms the
    whole scene.

    5 · Bright character portrait

    A cheerful anime adventurer poses in a crisp modern digital portrait with vibrant
    colours and soft skin shading. A waist-up framing angled slightly from the side catches
    her turning toward the camera, one hand raised in a small wave, the other steadying
    a satchel strap on her shoulder. Her round face shows flushed cheeks, bright teal eyes
    and a wide open smile. Her dark auburn hair is tied into two short braids with orange
    ribbons. She wears a rust-coloured travelling cloak over a cream tunic, leather bracers
    and a compass on a cord around her neck. Autumn woodland blurs softly behind her.
    Golden late-day light catches the loose strands of her hair.

    🔞 Adult-content scope

    In my experience with the original Genesis release, artistic nudity and adult pin-up imagery worked without an additional LoRA. Explicit sexual content was not the focus. No NSFW LoRA is baked into this checkpoint. V2’s changes described here concern training, detail, prompting and visual styles; they are not a new claim about adult-content capabilities.


    🔧 Installation

    Use the standard Krea 2 three-file setup. Download one diffusion-model variant and place the files in these folders:

    Diffusion model
    ComfyUI/models/diffusion_models/See-Krea2-Genesisv2-[format].safetensors

    Text encoder
    ComfyUI/models/text_encoders/qwen3vl_4b_bf16.safetensors

    VAE
    ComfyUI/models/vae/qwen_image_vae.safetensors

    For the BF16 master, the filename is See-Krea2-Genesisv2.safetensors, without a format suffix.

    1. Load Diffusion Model: choose your downloaded Genesis V2 file.

    2. Load CLIP: choose qwen3vl_4b_bf16.safetensors, type krea2.

    3. Load VAE: choose qwen_image_vae.safetensors.

    If you already run Krea 2 Turbo, reuse your existing text encoder and VAE. For quantized files, make sure your installed loader supports the selected format. Existing Krea 2 LoRAs can be tested with this checkpoint; their visual effect and compatibility still depend on the individual LoRA and workflow.


    📈 Version history

    V2 · More training, more detail, more styles

    • Longer training at higher resolution.

    • More image detail and improved prompt following.

    • Detailed anime women and detailed backgrounds in the training data.

    • English and Japanese text, plus manga material; improved text rendering remains unconfirmed.

    • New trained triggers: seesee_elf and seesee_kitsune.

    • Personal pixel-art look, prompted with pixel or 16Bit.

    • Game Boy-inspired direction trained with my own captures.

    • BF16 master and five compressed formats: INT8 ConvRot, FP8 Scaled, NVFP4, INT4 ConvRot and MXFP8.

    V1 · Initial Genesis release

    Anime-focused checkpoint on Krea 2 Turbo with an 8-step workflow at CFG 1.0. The original compressed builds included INT8 ConvRot, plain INT8, FP8 Scaled, NVFP4 and INT4 ConvRot. Plain INT8 belongs to the V1 lineup; it is not one of the V2 files listed above.


    🙏 Credits

    Base model: Krea 2 Turbo by Krea
    Text encoder: Qwen3-VL-4B
    VAE: Qwen-Image VAE
    Checkpoint and additional training: SeeSee


    📜 License

    This is a modified version of the Krea 2 model. It is not an official Krea product and is not endorsed by Krea.

    Use is subject to the Krea 2 Community License Agreement. Please read the applicable terms at Krea’s licensing page, including any conditions that apply to derivatives and commercial use.

    My additional permissions, subject to the base-model license:
    Commercial use: yes · Merging and derivatives: yes · Selling generated images: yes · Credit: appreciated.

    These permissions do not override the base-model license. If Genesis contributes to your own model, a mention of where the anime part came from is appreciated.

    Made by SeeSee · Built on Krea 2 Turbo · Expanded in V2 🎨

    Description

    Standard Text Encoders from Original Model

    FAQ

    Comments (8)

    seawolf338Aug 11, 2026
    CivitAI

    Prompt.

    aa9739614Aug 11, 2026· 1 reaction
    CivitAI

    各位大佬,506ti 16G,建议使用哪个版本,

    SeeSeeLP
    Author
    Aug 11, 2026

    BF16 and int8

    qingyun6663Aug 28, 2026

    @SeeSeeLP 
    16G显存可以用BF16?

    SeeSeeLP
    Author
    Aug 28, 2026

    @qingyun6663 yes 👍

    qingyun6663Aug 29, 2026

    @SeeSeeLP 
    另外询问下是不是8G版本的,就暂时不要考虑Int8及以上版本了?

    SeeSeeLP
    Author
    Aug 29, 2026

    如果你说的是 RTX 5060 Ti 8GB,那情况和 16GB 版本会有一些不同。

    并不是说 8GB 显存就“完全不能用 INT8 或更高精度版本”,而是要区分 能不能运行 和 是否值得使用。

    我的几个版本大致是:

    Genesis-int8_convrot:13.2GB

    权重误差约 0.99%

    目前几个量化版本里质量最好

    如果显存足够,我通常最推荐这个版本。

    Genesis-fp8_scaled:13.1GB

    权重误差约 2.66%

    比较适合 RTX 40 系及更新显卡。

    Genesis-int8:13.2GB

    权重误差约 1.55%

    和 int8_convrot 大小差不多,但更偏向推理速度。

    Genesis-nvfp4:7.7GB

    权重误差约 9.41%

    专门适合 RTX 50 系 / Blackwell。

    对 5060 Ti 8GB 来说,我个人首先建议测试这个版本。

    Genesis-int4_convrot:6.9GB

    权重误差约 16.43%

    显存占用最低,最适合低显存环境。

    如果 NVFP4 在你的工作流里显存还是比较紧,可以再试这个。

    所以如果你用的是 5060 Ti 8GB,我的推荐顺序会是:

    NVFP4 → INT4 ConvRot → INT8 / FP8(如果你愿意使用 RAM offload)

    尤其是 5060 Ti 属于 Blackwell 架构,所以 NVFP4 对这张卡很有意义。它的模型文件只有大约 7.7GB,比 13GB 左右的 INT8/FP8 小很多,而且可以利用 Blackwell 对 NVFP4 的支持。

    不过这里还有一个重要区别:

    模型文件大小不等于最终显存占用。

    除了 diffusion model 本身之外,ComfyUI 还需要加载:

    Qwen3-VL-4B text encoder

    Qwen Image VAE

    latent / attention 中间数据

    生成图片所需要的额外显存

    所以即使 NVFP4 文件是 7.7GB,也不代表它运行时只需要 7.7GB VRAM。

    但是 ComfyUI 可以把部分模型卸载到系统 RAM,所以 8GB 显卡并不是完全不能运行 13GB 的 INT8/FP8 模型。如果你有足够的系统内存,INT8 也有可能运行,只是会发生更多 CPU/RAM ↔ GPU 的数据交换,因此通常会明显慢一些。

    也就是说:

    8GB + INT8 = 可以尝试,但不是最理想的组合。

    8GB + NVFP4 = 我认为是 5060 Ti 8GB 最合理的选择。

    如果你的重点是“尽量保持最高质量”,你也可以测试 int8_convrot,让 ComfyUI 使用 RAM offload。它的权重误差只有 0.99%,非常接近 BF16 master。

    如果你的重点是“速度、显存占用和稳定性”,那我会直接选择 nvfp4。

    如果 NVFP4 在你使用的分辨率下仍然爆显存,再退到 int4_convrot。

    另外,不要太担心我页面里写的 9.41% 或 16.43% weight error。这个数字描述的是 量化后的模型权重与 BF16 原始权重之间的数学偏差,并不是说生成图片会“差 9%”或者“差 16%”。

    我实际做过几个版本的并排测试,生成结果之间的差异远远没有这个百分比看起来那么大。很多情况下必须把两张图放在一起仔细比较才能发现区别。

    所以对于 8GB 显卡,我反而建议:

    先实际试 NVFP4,不要因为 9.41% 的数字就直接认为画质会很差。

    简单总结:

    RTX 5060 Ti 16GB:
    int8_convrot 是我的首选,追求速度也可以用 int8,NVFP4 当然也可以。

    RTX 5060 Ti 8GB:
    我首先推荐 nvfp4。

    显存仍然不够的话:
    int4_convrot

    如果系统 RAM 很多,而且不介意速度下降:
    也可以尝试 int8_convrot / int8 / fp8_scaled,让 ComfyUI 做 offload。

    所以并不是“8GB 就不要考虑 INT8 以上版本”,更准确地说是:

    8GB 可以运行更大的版本,但 NVFP4 对 5060 Ti 8GB 来说通常是更合理、更均衡的选择。👍@qingyun6663 

    PrincesOfDarknessAug 15, 2026· 1 reaction
    CivitAI

    I really like this checkpoint; it does a good job.

    and offers useful tips—if I write the prompt well, I get exactly what I described.

    Checkpoint
    Krea 2

    Details

    Downloads
    159
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/11/2026
    Updated
    9/28/2026
    Deleted
    -

    Files

    seeKrea2_textEncoders_fp8.safetensors

    Mirrors

    HuggingFace (72 mirrors)
    ModelScope (1 mirrors)