CivArchive
    MiniMax H3_SparseRef15_Hybrid - Pruned_INT8
    NSFW

    【MiniMax H3_SparseRef15_Hybrid】

    MiniMax H3_SparseRef15_Hybrid is an experimental of Fl2VA-based hybrid models family that incorporates selected strengths of Ref2VA without any LoRA merging. It is designed to improve reference consistency while preserving natural motion, scene flexibility, and long-form stability.

    Note: This is not a native Ref2V model. It is based on Fl2VA with an added "Ref" effect, so please use LoRAs designed for Fl2VA rather than Ref2V LoRAs.

    The core concept is a sparse Ref2VA influence applied through AdaLN across 15 main transformer blocks, rather than using Ref2VA uniformly throughout the model.

    This sparse structure is intended to retain more of FL2VA's motion freedom and scene behavior while strengthening character identity and reference consistency.

    In testing, the model has shown strong resistance to character drift across long multi-clip generations, including cases where the character changes direction, temporarily leaves a clear frontal view, or continues through many consecutive clips.

    Different versions may vary in model size, precision, quantization, speed, memory usage, visual quality, and reference strength.

    For reference, Pruned_INT8 is a standard, lightweight pruned model in which the full model's "time-conditioning" has been compressed to 8 dimensions and the large-scale Attention/MLP matrices have been converted to the INT8 ConvRot format.

    In contrast, the Pruned_Partial-INT8 series also converts large-scale Attention/MLP matrices to the INT8 ConvRot format but reconstructs and retains critical components—such as AdaLN and time-conditioning—at FP32 precision.

    <Hybrid_Pruned_INT8>

    A lightweight Pruned Partial-INT8 version of SparseRef15 Hybrid.

    Technical characteristics:

    - 15 sparsely distributed Ref2VA AdaLN blocks within the 50-block main transformer

    - Pruned H3 architecture

    - Partial INT8 quantization

    - Designed for lower memory usage and faster inference

    - Strong emphasis on character identity consistency and long-form stability

    In extended multi-clip testing, this version maintained character identity with very little visible drift, even over long sequences.

    On suitable hardware, it is considerably faster and lighter than the heavier non-INT8 variant while retaining good motion quality and overall visual coherence.

    <Hybrid_Pruned_Partial-INT8 Ver.1.0>

    A newer Pruned Partial-INT8 Hybrid variant focused on stronger reference consistency and long-form character stability.

    Technical characteristics:

    - Pruned H3 architecture

    - Partial INT8 quantization for reduced model footprint and faster inference

    - Hybrid FL2VA / Ref2VA structure

    - Reference behavior tuned for stronger character identity persistence

    - Designed to remain stable across long multi-clip generations

    Compared with "Hybrid_pruned_int8", this version shows noticeably stronger reference persistence.

    In long-form testing, character identity remained highly stable across a 21-clip sequence with very little visible drift, even through repeated changes in pose, direction, and clip transitions.

    The stronger reference behavior can sometimes reduce scene freedom or resist situations where the character is expected to remain fully hidden for an extended period. In return, this version is particularly well suited to long-form generations where character consistency is the highest priority.

    The Partial-INT8 structure also gives this version a much smaller model footprint and significantly faster inference than the heavier non-INT8 Pruned variant on supported hardware.

    <Hybrid_Pruned_BF16>

    This is a hybrid derived from BF16 that has undergone only minimal pruning.

    The hybrid method is the same as Pruned_INT8.

    Description

    FAQ

    Comments (61)

    S1LV3RC01NAug 30, 2026
    CivitAI

    Did you actually use the int8 convrot models for the merge?

    Aki7777777
    Author
    Aug 30, 2026

    Yes. I used minimax_h3_fl2va_pruned_int8_convrot as the base and minimax_h3_ref2va_pruned_int8_convrot as the overlay. The MiniMax H3 Hybrid Loader is quantization-aware and preserves the INT8 ConvRot / .comfy_quant structure when overlaying the selected AdaLN blocks.

    Aki7777777
    Author
    Sep 6, 2026

    @S1LV3RC01N
    By the way, the latest version is based on the full BF16 version.

    Aki7777777
    Author
    Aug 30, 2026· 33 reactions
    CivitAI

    Thanks for trying SparseRef15.
    This is an experimental distributed FL2VA / Ref2VA hybrid.
    I’d love to hear what works best for you: short clips, chained video, stronger reference use, or something else.

    vAnN47Aug 30, 2026· 1 reaction

    testing it now, hope i'll get good results, uploading soon some pdd lora tests vs dareties lora with the hybrid checkpoint 30-49. would like really see the diff with your checkpoint

    Aki7777777
    Author
    Aug 30, 2026· 2 reactions

    @vAnN47 
    Personally, I recommend LightX2V Turbo 4-step v0.1.

    kaifesalazar431Sep 3, 2026

    Would you recommend a base workflow to start with? :)

    Aki7777777
    Author
    Sep 3, 2026· 1 reaction

    @kaifesalazar431 
    I’d recommend starting with a simple workflow built only with nodes included in the latest version of ComfyUI, such as UNETLoader, MiniMaxH3ImageToVideo, BasicGuider, BasicScheduler, and SamplerCustomAdvanced.

    I’d suggest adding optional nodes such as ModelSamplingMiniMaxH3 only later, once you have a basic workflow working well.

    Also, H3 low-step LoRAs can behave surprisingly differently, so it’s definitely worth trying a few. :)

    larflarfSep 7, 2026· 1 reaction

    The pruned partial version references the original image much better :D

    vAnN47Sep 7, 2026· 4 reactions

    @larflarf interesting, worth downloading? im started to use mods (instead of char lora) wonder how it will behave.

    larflarfSep 7, 2026· 2 reactions

    @vAnN47 Of course. The pruned partial version seems better for use everywhere. In fl2v, i2v, and ref2v alike. 😀

    Aki7777777
    Author
    Sep 8, 2026· 4 reactions

    @vAnN47 @larflarf
    Pruned_INT8 Ver.0.1 is a standard lightweight pruned model that compresses the full model's time-conditioning to 8 dimensions and converts large Attention/MLP matrices into INT8 ConvRot format.

    In contrast, Pruned_Partial-INT8 Ver1.0 also converts large Attention/MLP matrices to INT8 ConvRot format but retains critical components—specifically AdaLN and time-conditioning—reconstructed at FP32 precision.

    koongrizzSep 19, 2026

    i use it like this : your checkpoint + turbo v4 step600 ema pruned lora at 10 steps + spectrum forecast on, gives best results, still fast, no longer using sage/sol/sla attention because it changed the results to much when i tested with reference images and reference sound.

    MikeflowerAug 30, 2026· 8 reactions
    CivitAI

    Your model is not NSFW. Your examples are incorrect🫤

    Aki7777777
    Author
    Aug 30, 2026

    Should I not have used it to mean that it is capable of generating NSFW content?

    g1263495582Aug 30, 2026

    For this one, I think your prompting skills might not be up to par.

    Aki7777777
    Author
    Aug 30, 2026

    @g1263495582 
    This is not a model exclusively for NSFW content. I apologize if I caused any misunderstanding.

    DaddyWolfgangAug 31, 2026· 1 reaction

    Uh, MiniMaxH3 by default is NSFW. No need to apologize @Aki7777777 because it's obvious this person is trying to sow discord.

    kingdan78Aug 30, 2026· 6 reactions
    CivitAI

    Good job! But very Asian oriented ;) all my caucasian girls end asian

    Aki7777777
    Author
    Aug 30, 2026· 1 reaction

    Thanks! Though, I haven't really made any changes to the default settings...

    Aki7777777
    Author
    Aug 31, 2026· 1 reaction

    I tried it on non-Asians, too. What do you think?

    kingdan78Aug 31, 2026· 1 reaction

    @Aki7777777 wow looking real good... makes me wonder why my references end up looking asian hahaha good job :)

    mwoody450Aug 31, 2026· 1 reaction

    It's not your imagination: it definitely converted my caucasian subjects to Asian, especially once they were far enough from the camera.

    Aki7777777
    Author
    Sep 1, 2026

    @mwoody450 @kingdan78 
    Interesting. Since this hybrid is essentially based on minimax_h3_ref2va_pruned_int8_convrot, the tendency may originate from the original Ref2VA model. I haven't seen it in my own generations, so it may be more noticeable with extremely long single-clip generations.

    Aki7777777
    Author
    Sep 6, 2026

    @mwoody450  @kingdan78 
    A phenomenon has been observed where the character's appearance changes due to the use of low-step LoRA.

    I recommend trying out multiple LoRAs.

    velantegAug 31, 2026· 3 reactions
    CivitAI

    Regardless of quality if model can draw only asians its auto skip.

    Aki7777777
    Author
    Aug 31, 2026

    I have simply combined off-the-shelf models; no intentional customization has been performed.

    DaddyWolfgangAug 31, 2026· 1 reaction

    Did it ever occur to you that you can prompt for other races? You should try it.

    kkmw15Aug 31, 2026· 2 reactions
    CivitAI

    It is a very excellent model. Thank you.

    transformermanAug 31, 2026· 4 reactions
    CivitAI

    Well, this replaced the regular ref2va for me! It's not even close, I much prefer this one. Movement is more natural. Physics make more sense. Prompts are better followed. It seems to let me use ref images for styling, that don't immediately become part of the events. Very nice!

    So far, I've just used it like the regular ref2va model. I do prefer the 4step_v1.1_768p turbo lora with it.

    Thank you for your efforts!

    Aki7777777
    Author
    Aug 31, 2026

    Thank you so much for the detailed feedback!
    This is very close to what I hoped the sparse hybrid structure might achieve — keeping useful Ref2VA conditioning while giving FL2VA more freedom for motion and prompt interpretation.

    Your note about using reference images mainly for styling without having them immediately become part of the events is especially interesting. I'll definitely keep an eye on whether other users observe the same behavior.

    Also, thanks for the 4step_v1.1_768p Turbo LoRA recommendation. I haven't tested that combination enough yet, so I'll give it a try!

    Pat3dxSep 1, 2026

    @Aki7777777 the new turbo lora v1.1 4 step 768 is for FL2VA, not R2V, and it not work good for R2V..
    If you want production ready quality for your R2V video, just use the turbo v4 step600 ema hybrid lora, that work for both model. with at least 8 steps.

    Aki7777777
    Author
    Sep 1, 2026

    @Pat3dx 
    Thank you for the advice.

    As stated in the description, this model is based on FL2VA, so it is perfectly fine to use the FL2VA Fast LoRA with it.

    big27916430Sep 2, 2026· 1 reaction

    @Aki7777777 Yes, in the short term, such as 5S, it may be executed immediately, but in the long term, such as 10s, there will be a process from the beginning to the reference posture, but this can be controlled through prompt words.

    The characters set will reference the poses in the pose map in the specified environment, instead of completely replicating the pose map like the original version.

    Aki7777777
    Author
    Sep 2, 2026

    @big27916430 
    Thanks a lot for the detailed feedback!
    I'm really glad to hear the timing-based prompt control is working well for you, especially for transitions toward the reference pose in longer clips.
    It’s also great to hear that the imitative action behavior is stronger than in the original version.
    I really appreciate your testing and support!

    OrangeJuiceAlienSep 1, 2026· 3 reactions
    CivitAI

    I see the fine detail quality and video consistency is much better than with pure ref2va model. unfortunately reference following is noticeable worse than pure ref2va model. so I would wish for a version that tries to keep more ref following intact, even maybe at small cost to quality.

    Aki7777777
    Author
    Sep 1, 2026· 5 reactions

    Thanks for the feedback!
    I'm actually working on a version that keeps stronger Ref2VA reference following, even if it comes at a small cost to detail quality or consistency. I'm still testing the balance, but I'll share it once it's ready.

    OrangeJuiceAlienSep 2, 2026· 1 reaction

    @Aki7777777 I will wait for it!

    Aki7777777
    Author
    Sep 6, 2026

    @OrangeJuiceAlien
    I have confirmed that "Pruned Partial-INT8 V1.0" exhibits a strong "ref" effect.

    I'm not sure if it will suit your preferences, but please give it a try.

    OrangeJuiceAlienSep 12, 2026

    @Aki7777777 sadly I still see the new version has much worse reference following than the pure ref model

    Aki7777777
    Author
    Sep 13, 2026

    @OrangeJuiceAlien 
    Yes, that is expected. The pure Ref model is designed to maximize reference following, while this version intentionally trades some reference strength for better motion freedom, scene behavior, and overall balance.
    so it is not intended to outperform the pure Ref model in reference strength alone.

    mangho9123389Sep 1, 2026
    CivitAI

    That's really cool. It's fascinating. Does this model represent the genital area better than the stock base model?

    Aki7777777
    Author
    Sep 1, 2026

    Thank you for your feedback.

    Essentially, it is based on FL2v, with Ref2VA AdaLN layers distributed and referenced at intervals throughout the DiT.

    If you are looking for even more realistic rendering or movement, I recommend using FL2v-based LoRAs.

    baoanhnguyenkts784Sep 1, 2026· 3 reactions
    CivitAI

    been testing yours and eros3, yours is better overall when using FL2AV for ref workflow, the trade-off is to sacrifiy the first frame which nothing to me, Thank you.
    Just minor suggestion, may you check this technique out? https://huggingface.co/cicalooo/10Eros-Max-h3-int8-convrot/blob/main/SKIP_EDGES.md they claim the the first and last transformer blocks would be fully taken in account.

    Aki7777777
    Author
    Sep 1, 2026

    Thanks! I'm glad to hear it's working well for your FL2AV reference workflow.
    I checked the skip-edges technique. It looks like it keeps blocks 0, 1, 48 and 49 in BF16 while quantizing the rest, rather than directly changing the reference weighting.
    My current hybrid is built from already-quantized INT8 ConvRot checkpoints, so applying this properly would require rebuilding it from the BF16 source weights.
    It's an interesting idea though, so I'll keep it in mind for a future test.

    velantegSep 7, 2026

    Did you used 20 steps? Any turbo loras give bad result with this model, Eros much better.

    Aki7777777
    Author
    Sep 7, 2026

    @velanteg 
    If you're using a low-rank LoRA, 20 steps shouldn't be necessary. Which LoRA are you using?

    big27916430Sep 1, 2026· 5 reactions
    CivitAI

    Testing complete. This is undoubtedly the best ref2va model I've tested! It features excellent prompt adherence and reference consistency, with outstanding image clarity.

    One observation from my tests: applying LoRAs to other models often leads to severe structural distortions (e.g., human anatomy or perspective issues). For future updates, I highly recommend keeping the base model clean and letting users apply LoRAs on their own, rather than merging them by default.

    Really appreciate your hard work and contribution to the community!

    (P.S. Translated from AI, so please excuse any slight translation artifacts.)

    Aki7777777
    Author
    Sep 1, 2026

    Thanks for the detailed feedback!
    No worries — this model does not have any LoRAs merged into it. It is built only by combining selected weights from the original FL2VA and Ref2VA models.
    I also prefer to keep LoRAs separate so users can apply them freely depending on their workflow.

    velantegSep 3, 2026

    Are you joking? I cant make any not messed video with this model.

    Aki7777777
    Author
    Sep 3, 2026

    @velanteg 
    What kind of workflow are you using? You shouldn't have any issues if you use a standard workflow combined with the appropriate LoRA as needed. I would appreciate it if you could tell me in what way the generated videos are "not right."

    velantegSep 4, 2026

    @Aki7777777 2 stage workflow

    Aki7777777
    Author
    Sep 4, 2026

    @velanteg 
    What exactly do you mean by a "two-stage workflow"? Since H3 itself does not have distinct stages like "High" or "Low," it is unclear what kind of configuration you are referring to. If you are referring to a workflow where outputs from multiple samplers are concatenated at the final frame, I have confirmed that using 10 or more stages works without issues. (The video I posted of the running rabbit uses about 15 stages.)

    velantegSep 4, 2026

    @Aki7777777 low res + upscale.

    Aki7777777
    Author
    Sep 4, 2026

    @velanteg 
    Generating at low megapixels causes fine details to be lost, so I do not recommend it.

    Are you running out of VRAM?

    big27916430Sep 5, 2026· 1 reaction

    @velanteg From my comment, you can see that this is a 'raw model' without adding any Lora. So, a small number of sampling steps will not be able to complete the image structure, and you need more sampling deployment or adding tubro Lora

    Whistler42Sep 2, 2026
    CivitAI

    Would it be possible to make a bf16 version also?

    Aki7777777
    Author
    Sep 2, 2026

    BF16 version is definitely possible. I'm currently looking into higher-precision variants, although the file size will be quite large.

    tomgm777320Sep 6, 2026

    I would also like the bf16 type.

    Aki7777777
    Author
    Sep 6, 2026· 2 reactions

    @tomgm777320 
    With the pure BF16 format, offloading is effectively mandatory, even on an RTX 5090.

    Even the pruned version remains extremely "heavy," so it is unlikely to be practical for most users.

    I'm currently fine-tuning the model to reduce its size while minimizing any loss in performance, so please bear with us a little longer.

    Aki7777777
    Author
    Sep 8, 2026

    @Whistler42 @tomgm777320
    It runs heavily and hasn't been fully tested, but I have released the Prude version.

    I would love to hear your thoughts if you'd like to share them.

    Checkpoint
    MiniMax H3

    Details

    Downloads
    2,857
    Platform
    CivitAI
    Platform Status
    Available
    Created
    8/30/2026
    Updated
    9/20/2026
    Deleted
    -

    Files

    minimaxH3Sparseref15_prunedINT8.safetensors