CivArchive
    MiniMax H3 Continuum – Long-Form Video & Audio for ComfyUI - v1.0
    NSFW
    Preview 139784881
    Preview 139755078
    Preview 139760008

    MiniMax H3 Continuum

    GitHub:
    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    [ V3.5 ]

    V3.5 introduces two major additions:

    - Continuum-aware Second Pass / Hi-Res Fix

    Refine externally processed H3 latents while preserving Continuum physical groups, prompts, seeds, ordering, and first-pass audio. An integrated one-node 2x Hi-Res Fix path is also included as an experimental feature.

    - Low-memory Assemble + Seam V3.5

    Adds Auto, RAM, and Disk-backed video-buffer modes. Disk-backed assembly significantly reduces system RAM/private-memory usage for long or high-resolution outputs while preserving Exact Duration, Seam, Terminal Merge, and audio behavior.

    All V3.4 nodes remain available for saved-workflow compatibility. Existing V3.4 workflows continue to work unchanged.

    The V3.5 release passed 430 automated tests and representative GPU acceptance tests.

    Note: The integrated Hi-Res Fix remains experimental. Long 2x workflows can require substantial GPU VRAM.

    #5 3 chunks x 15 seconds

    #6 6 chunks x 15 seconds

    W576xH576

    [ v3.4 ]

    Long-form MiniMax H3 video and audio generation for ComfyUI with chunked generation, persistent references, restartable runs, and user-controlled audio.

    ### What's new in v3.4

    - Driving Audio: preserves the supplied audio as the final audio while guiding generation across chunks.

    - Video Reference: provides persistent visual reference for identity, motion, framing, and scene appearance.

    - Restartable chunks: reuse completed chunks with Run Storage and regenerate only the required part.

    - Improved Core compatibility: unknown upstream or custom nodes are not rejected merely because they are not recognized by Continuum.

    - Simpler stable interface: obsolete compatibility controls and experimental Timeline inputs are hidden from the V3.4 public workflow.

    - Spectrum interoperability: Spectrum remains optional and can use the official H3 Continuum Interop API.

    ### Direction change from v3.3

    V3.4 focuses on predictable reference workflows rather than experimental Timeline Video and timeline-audio generation.

    Driving Audio preserves the original user-supplied audio. Video Reference provides persistent visual guidance without requiring exact frame-by-frame copying. Existing V3.3 workflows remain available through legacy compatibility paths.

    ### Updating

    For an existing Git installation:

    git pull --ff-only origin main

    V3.4 input connection patterns

    V3.4 separates the visual reference input from the driving-audio input. Choose the connection pattern that matches your source material.

    1. Audio only

    Connect Load Audio to driving_audio. Use this when an existing song, dialogue track, or sound effect should remain the final audio. A Video Reference is not required.

    Driving Audio connection

    2. Video with its own audio

    Connect Load Video (Upload) IMAGE to Video Reference. If the uploaded video contains the audio you want to preserve, connect its AUDIO output to driving_audio as well.

    Video Reference and embedded audio connection

    3. Video and audio from separate sources

    Connect Load Video (Upload) IMAGE to Video Reference, then connect a separate Load Audio node to driving_audio. Use this when the visual reference video and the final audio source are different files.

    Separate Video Reference and Driving Audio connection

    Both inputs are optional. Connect Video Reference when visual guidance is needed, and connect driving_audio when the supplied audio should be preserved in the final output.

    Video Reference frame rate

    Use a 24 fps source for Video Reference. Load Video (Upload) may accept files recorded at 25 fps or another frame rate, but acceptance alone does not guarantee correct temporal alignment with H3. For a non-24 fps source, set force_rate to 24 in Load Video (Upload), or convert the file to 24 fps before loading it. If the source is already 24 fps, leave force_rate at its default and do not resample it.

    Current validation status

    [ v3.3 ]

    V3.3 adds Timeline Video conditioning for long-form MiniMax H3 generation. A reference video can now be processed in chunk-local time slices, allowing motion and scene continuity to be carried across multiple 5-second chunks while keeping the reference resolution independent from the output resolution. The Efficient 0.4 MP mode helps reduce memory usage and processing time.

    Video assembly has also been improved. Auto seam handling analyzes chunk boundaries and applies guarded corrections for transient flicker, micro-flash, exposure, and color differences. This helps produce more natural transitions between generated chunks without changing the original sampling process.

    Existing V3.2.4 workflows remain available as Legacy nodes for compatibility.

    [ v3.24 ]


    Generate longer native MiniMax H3 video and audio sequences in ComfyUI.

    H3 Continuum is a ComfyUI custom node that generates a longer sequence as connected chunks and assembles them into one continuous video.

    ```text
    3 × 5-second chunks → 15-second video
    6 × 5-second chunks → 30-second video

    The previous video and audio latent context is passed into each continuation chunk. This is not a simple video concatenation workflow.

    Main purpose: longer MiniMax H3 generation, not faster generation.

    Easy Installation

    H3 Continuum can be installed directly from ComfyUI Manager.

    1. Open ComfyUI Manager

    2. Search for H3 Continuum or Continuum

    3. Select Install

    4. Restart ComfyUI

    5. Load one of the included sample workflows

    Manual installation and the latest documentation are available on GitHub:

    GitHub:
    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    What It Does

    H3 Continuum divides a longer generation into manageable chunks.

    MiniMax H3 Model
    ↓
    H3 Continuum Sampler
    ↓
    ComfyUI Core Video / Audio VAE Decode
    ↓
    H3 Continuum Assemble
    ↓
    Final video

    Each continuation chunk receives latent context from the preceding chunk. Overlapping context is removed during assembly, and the final frame and audio counts are aligned to the requested duration.

    Main Features

    • Connected long-form MiniMax H3 generation

    • Native video and audio latent continuation

    • Fixed, List, and Timeline prompt formats

    • Automatic prompt-format detection

    • T2VA, I2VA, FL2VA, Last Frame and Reference workflows

    • Up to three Reference Images

    • Reference Audio conditioning

    • First Frame and Last Frame conditioning

    • Configurable continuity context

    • Run Storage and automatic resume

    • Partial regeneration from a selected chunk

    • Optional Spectrum interoperability

    • Standard and Turbo sample workflows

    • ComfyUI Core VAE Decode compatibility

    Included Sample Workflows

    Two example workflows are provided.

    Standard Workflow

    Recommended when output quality and temporal consistency are the priority.

    • Standard MiniMax H3 sampling

    • Spectrum can be enabled

    • Suitable for quality-focused generation

    • Reference Image and Reference Audio supported

    • RTX upscaling can be enabled when required

    Turbo Workflow

    Recommended for faster tests and iteration.

    • LightX2V MiniMax H3 Turbo LoRA

    • 8-step example configuration

    • Spectrum is bypassed by default

    • Faster than the standard workflow in tested configurations

    • Some loss of facial detail or additional artifacts may occur

    Turbo LoRA models:

    https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main

    MiniMax H3 models and documentation:

    https://huggingface.co/MiniMaxAI/MiniMax-H3

    Models and LoRAs are not included with this custom node.

    Reference + Continuation

    Reference Images remain available across all generated chunks.

    A typical setup is:

    Picture 1 → face and identity
    Picture 2 → full-body appearance and clothing
    Picture 3 → environment or an additional visual reference
    Audio 1   → vocal, music or audio-performance reference

    Ref2VA is the reference-specialized checkpoint and is generally the first choice for stronger reference fidelity.

    FL2VA with Reference conditioning is also allowed. H3 Continuum does not automatically replace or switch the connected model.

    Spectrum Integration

    Spectrum is optional. H3 Continuum also works without it.

    With a compatible Spectrum release, H3 Continuum sends a continuation signal only when generating later chunks.

    Chunk 1 → normal Spectrum sampling
    Chunk 2+ → Continuum Actual Prefix 2

    This allows Spectrum to coordinate its spectral forecasting with the continuation context instead of treating every chunk as an unrelated generation.

    Benefits include:

    • Automatic identification of continuation chunks

    • Actual Prefix applied only where required

    • No manual prefix switching between chunks

    • Reduced risk of duplicated prefix processing

    • Compatibility with standard ComfyUI workflow execution

    Spectrum remains an approximate accelerator. Motion, anatomy, audio and detail can differ from a non-Spectrum result, so quality comparisons should use the same prompt and seed.

    Spectrum:

    https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

    Run Storage and Resume

    Enable Save + Auto Resume to preserve completed raw video and audio chunks.

    If a generation is interrupted, H3 Continuum can reuse compatible saved chunks and continue from the first missing chunk.

    It can also regenerate from a selected chunk while preserving the compatible prefix.

    Chunk 1–3 completed
    ↓
    Generation interrupted
    ↓
    Queue the workflow again
    ↓
    Chunks 1–3 reused
    ↓
    Generation continues from Chunk 4

    Run Storage verifies the sampling contract, model route, references, resolution and saved chunk files before reuse.

    Prompt Formats

    Fixed

    One prompt is used for every chunk.

    List

    Separate prompts are divided with:

    ---

    Timeline

    [0-5s]
    First scene description
    
    [5-10s]
    Second scene description
    
    [10-15s]
    Third scene description

    Prompt Format = Auto detects the appropriate format automatically.

    Incomplete timeline coverage produces diagnostics and safe fallback behavior rather than unnecessarily stopping every generation. Structurally unusable input is still reported as an error.

    Tested Configuration

    The current Windows implementation has been tested with:

    GPU                 NVIDIA RTX 5060 Ti 16GB
    ComfyUI             MiniMax H3-compatible Core build
    Chunk Duration      5 seconds
    Typical Length      3 or 6 chunks
    Continuity          Balanced 22 frames
    Standard Sampling   RES Multistep
    Spectrum Interop    Actual Prefix 2

    The node is not limited to RTX 50-series GPUs. Actual compatibility, generation speed and usable resolution depend on the MiniMax H3 model, GPU memory, ComfyUI configuration and installed acceleration nodes.

    RTX 4060 and other configurations have not been formally validated by this project.

    Frequently Asked Questions

    Is this only a workflow?

    No. H3 Continuum is a ComfyUI custom node package. The included workflows are ready-to-use examples.

    Does it generate one native 30-second sample?

    No. It generates connected chunks and assembles them into one longer output while carrying video and audio latent context forward.

    Does it make MiniMax H3 faster?

    Speed is not the primary purpose. H3 Continuum is designed for longer generation. Spectrum and Turbo LoRAs can reduce generation time in some configurations.

    Is Spectrum required?

    No. It is an optional acceleration and interoperability path.

    Can I use the Turbo LoRA?

    Yes. A Turbo sample workflow is provided. Spectrum is bypassed by default in that workflow because combining both can change quality or introduce artifacts.

    Which model should I use for Reference Images?

    Ref2VA is the reference-specialized option. FL2VA with Reference conditioning is also allowed, but reference fidelity may differ.

    Are the models included?

    No. MiniMax H3 checkpoints, text encoders, VAEs, Turbo LoRAs and optional acceleration nodes must be installed separately.

    Can an interrupted generation be resumed?

    Yes. Enable Run Storage before generation. Compatible completed chunks can then be reused.

    Can I regenerate only the later part?

    Yes. Run Storage supports regeneration from a selected chunk while retaining a compatible earlier prefix.

    Are chunk boundaries always invisible?

    No generative continuation system can guarantee a completely invisible boundary. H3 Continuum preserves latent context and removes duplicated overlap, but difficult motion, lighting changes and large prompt transitions can still produce flicker or visual changes.

    Does Reference Audio guarantee exact lip synchronization?

    Reference Audio conditions MiniMax H3’s native joint video/audio generation. It can guide vocals, rhythm, expression and mouth movement, but it does not guarantee sample-identical audio reproduction or frame-perfect lip synchronization in every generation.

    Does it support audio continuity?

    Yes. Video and audio latent context are carried together. The assembler also provides an optional Audio Seam mode for boundary-local audio correction.

    Is RTX 5090 required?

    No. Development and runtime validation were performed on an RTX 5060 Ti 16GB. Lower-memory configurations may require reduced resolution, offloading or other ComfyUI memory optimizations.

    What license is used?

    H3 Continuum is released under the MIT License.

    Custom Nodes

    The included Standard and Turbo workflows use the following custom nodes.

    - H3 Continuum

    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    - rgthree-comfy

    https://github.com/rgthree/rgthree-comfy

    - ComfyUI-Easy-Use

    https://github.com/yolain/ComfyUI-Easy-Use

    - ComfyUI-KJNodes

    https://github.com/kijai/ComfyUI-KJNodes

    - ComfyUI-Spectrum-MiniMax-H3

    https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

    - NVIDIA RTX Nodes for ComfyUI

    https://github.com/Comfy-Org/Nvidia_RTX_Nodes_ComfyUI

    Spectrum and RTX upscaling are optional generation paths, but installing all listed custom nodes allows the included workflows to load without missing-node warnings.

    Models

    - MiniMax H3

    https://huggingface.co/MiniMaxAI/MiniMax-H3

    - LightX2V MiniMax H3 Turbo LoRA

    https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main

    Models and LoRAs are not included in the workflow ZIP.

    Main Links

    - GitHub and documentation

    https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum

    - Install from ComfyUI Manager

    Search for H3 Continuum

    Description

    Initial public release. Includes Standard and Turbo workflows for H3 Continuum V3.2.4.

    FAQ

    Comments (42)

    greggyq159Aug 15, 2026· 3 reactions
    CivitAI

    My Dude! Thank you so much! This works amazingly. Better than I could have hoped. I was even able to integrate the nodes into my personal workflows no problem.

    I have one question though. A frustration I've had with using reference videos in the reference to video workflow is that if the video is more than a few seconds, or of high resolution, vram blows up. And everything crashes and burns. I've done little tricks with video editors, cutting things into clips and lowering resolution. But it's a pain in the butt. It would be awesome if you could "chunk" a video to video edit utilizing nodes like this.

    greggyq159Aug 16, 2026

    I did actually realize I have a problem. Maybe I missed something somewhere. Can't figure out how to restart a generation on a later chunk? Like if I like the first chunk but not the 2nd chunk. I don't want to have to generate the 1st one again, I want it to start from producing the 2nd chunk? How do I do that?

    ukr8b3g201
    Author
    Aug 16, 2026· 2 reactions

    @greggyq159 For restarting from a later chunk in the current V3.2.4 workflow:

    1. Set:

    Run Storage = Save + Auto Resume

    Run Name = blank, unless you want to use a manual name

    Regenerate From = Chunk 2

    2. Queue the same workflow again.

    The automatic Run Storage ID is generated from the sampler configuration when Run Name is blank. You do not normally need to create or enter an ID manually.

    With Regenerate From set to Chunk 2, the saved chunks before Chunk 2 are reused, and Chunk 2 and later chunks are generated again.

    Keep the model, reference inputs, resolution, continuity, and other contract-related settings compatible with the saved run. Changing those settings may create a different Revision and prevent reuse.

    Run Storage must be set to Save + Auto Resume; Regenerate From does not work as a reuse operation when storage is disabled.

    123sirako123521Aug 15, 2026· 2 reactions
    CivitAI

    A little confused on the prompting. For Ref2V, do I have to define the subjects, shot summary and scenes 3 seperate times? Or write 1 long 30 second prompt and set it to 6 5 second chunks. Could you share some examples?

    ukr8b3g201
    Author
    Aug 16, 2026· 1 reaction

    You do not need to redefine the subjects three separate times.

    For Ref2VA, define the reference relationships once:

    - Subject definitions

    - Summary

    - Retention analysis

    - Detailed description

    - Overall soundscape

    - Non-diegetic music

    The reference images and optional reference audio remain available throughout the generation.

    For H3 Continuum, however, the visual timeline should be divided into one section per 5-second chunk. For a 30-second video, use six timeline sections rather than one unstructured 30-second paragraph.

    For example:

    [0-5s]

    The woman stands in the opening composition and begins walking slowly toward the camera.

    ---

    [5-10s]

    She continues walking along the same street while the camera tracks backward.

    ---

    [10-15s]

    The camera moves slightly to her side as she looks toward the storefronts.

    ---

    [15-20s]

    She slows down and turns naturally toward the camera.

    ---

    [20-25s]

    The camera makes a gradual push-in while she continues the same action.

    ---

    [25-30s]

    She reaches the end of the movement and holds the final pose.

    The complete Ref2VA reference definitions are written once. Only the action, camera movement, and scene progression change from chunk to chunk.

    In practice:

    - Standard Ref2VA: one short video generation with persistent references.

    - H3 Continuum: multiple 5-second generation chunks with persistent references and separate timeline instructions.

    Use --- separators or explicit time ranges so Continuum can detect the individual chunk prompts. Avoid repeating the entire subject description six times, and avoid using one long unstructured paragraph for a multi-chunk generation.

    ukr8b3g201
    Author
    Aug 16, 2026· 1 reaction

    Clarify that:

    1. In Ref2VA, subject definitions, summary, retention analysis, and other reference relationships are written once.

    2. Reference Images and Reference Audio remain persistent across Continuum chunks.

    3. For Continuum, the visual timeline should be divided into one prompt section per 5-second chunk.

    4. Do not repeat the complete subject definitions for every chunk.

    5. Do not rely on one unstructured 30-second paragraph. Use six clearly separated 5-second sections, with --- separators or explicit time ranges such as [0-5s], [5-10s], and so on.

    6. Each chunk should describe the action, camera movement, and scene progression for that interval.

    7. The prompt must match the number of requested chunks.

    8. Make clear that this is for long-form Continuum generation, not ordinary single-pass Ref2VA.

    Include a short six-chunk example using:

    - one woman as the main subject,

    - two reference images,

    - one optional reference audio,

    - six 5-second chunks,

    - consistent identity and clothing,

    - gradual camera and action progression.

    Do not claim guaranteed frame-perfect continuity or exact audio copying.

    nekotxtAug 16, 2026

    @ukr8b3g201  So we can basically use the 'standard' Minimax prompt format for Ref2V, with the only remark being that the shots should all be 5 seconds long? That would make the most sense for me.

    ukr8b3g201
    Author
    Aug 16, 2026

    @nekotxt Almost, but the 5-second rule applies to Continuum chunks, not individual shots.

    You can use the standard Ref2VA prompt structure. Only the timeline needs one prompt section for each generated chunk. With the current recommended setup, each chunk is 5 seconds:

    [0-5s]

    [5-10s]

    [10-15s]

    A 5-second chunk may contain one continuous shot or several shorter shots. They do not all have to be exactly five seconds long.

    For better continuity, however, it is usually safer to keep one main action or gradual camera progression within each chunk.

    nktmanyik69Aug 15, 2026· 3 reactions
    CivitAI

    This is exactly what I was looking for but it doesn't have the most important input; reference video

    ukr8b3g201
    Author
    Aug 16, 2026

    If you are not using Continuum's chunk-based long-form processing, I recommend using the standard ComfyUI H3 Reference-to-Video workflow instead.

    Timeline Video is mainly useful when you want to process a longer sequence in separate Continuum chunks while using a video as motion and composition guidance.

    For the experimental Timeline Video node, the reference video is processed at approximately 0.4 MP by default. Using Match Output can increase processing time and memory usage significantly, so the lower reference resolution is intentional. It is used as a lightweight visual reference, not as a direct video replacement.

    The purpose is not to reproduce the source person or video frame by frame. The generated subject may be different. The Timeline Video provides guidance for motion, camera movement, and scene composition, while Reference Images can be used separately for identity, appearance, and clothing.

    In short:

    - Standard Ref2VA: use this for ordinary image/audio reference generation.

    - Timeline Video + Continuum: use this for chunked long-form generation with motion and camera guidance.

    - Timeline Video does not provide exact video-to-video replacement or guaranteed identity preservation.

    The Timeline Video workflow is still experimental and is being tested locally for VRAM usage, quality, and long-sequence stability.

    RusMaltsevAug 16, 2026· 1 reaction
    CivitAI

    [ERROR] !!! Exception during processing !!! H3 runtime is incompatible: PackedLayout no longer accepts 'frame_count'

    ukr8b3g201
    Author
    Aug 16, 2026

    Could you please share the complete ComfyUI error traceback, starting from

    “!!! Exception during processing !!!” through the final exception line?

    The Preview Text status is not enough to identify the compatibility issue.

    RusMaltsevAug 16, 2026· 3 reactions

    @ukr8b3g201 , node: "H3 Continuum Sampler V3.2" > "strict_compatibility" - [false]! (default - true), and the problem disappeared.

    ukr8b3g201
    Author
    Aug 16, 2026

    @RusMaltsev Thank you for finding this.

    Setting strict_compatibility to false disables the strict checkpoint compatibility gate. This can allow experimental combinations such as FL2VA with Reference conditioning to run instead of stopping with a compatibility error.

    Please note that this is not a general fix for every error. It only relaxes the compatibility validation. For standard Ref2VA workflows, keeping strict_compatibility=true is safer. With it disabled, the node may proceed with an unverified model configuration, so reference fidelity is not guaranteed.

    If the error is:

    PackedLayout no longer accepts 'frame_count'

    that is a separate ComfyUI Core API compatibility issue and requires a runtime update or a compatible ComfyUI revision. The strict compatibility setting should not affect that error.

    Please share the complete traceback if that specific error still occurs.

    Complement_CascadeAug 18, 2026

    @ukr8b3g201 I also got this error and it appears to be related to an update to comfy which changed how it native H3/packelLayout is expected, rolling back to a previous comfy fixed the issue.

    ukr8b3g201
    Author
    Aug 18, 2026

    @Complement_Cascade This error is caused by a ComfyUI Core API change, not by the strict compatibility setting.

    Recent Core revisions changed the PackedLayout constructor and removed the frame_count argument. Please either update H3 Continuum to a version that supports your current ComfyUI revision, or temporarily use the ComfyUI revision that was compatible with the workflow.

    Setting strict_compatibility to false only bypasses the model compatibility check. It does not fix the PackedLayout API mismatch.

    Complement_CascadeAug 18, 2026

    @ukr8b3g201 Latest version of comfy and latest version of H3 Continuum. From what I could see it was an issue with a change to base comfy that changed PackedLayout API. and H3 Continuum was not yet updated to reflect these changes. So my understanding was in order to use this wonderful node (I really love it) the easiest way was to rollback comfy. (I am NOT going to complain that a node is not updated fast enough for my convenience as it is free and a really nice node). If I have misunderstood I take full blame -also I have not used the Setting strict_compatibility to false. that was the previous poster and I was just trying to share info on the actual issue and a dirty fix

    ukr8b3g201
    Author
    Aug 18, 2026· 1 reaction

    @Complement_Cascade You are correct, and there is no need to take any blame.

    I checked the current ComfyUI Core implementation. PackedLayout now accepts:

    text_len, latent_t, latent_h, latent_w, audio_t, keyframes, and refs

    The old frame_count constructor argument has been removed. Therefore, with the latest published H3 Continuum version at the time, rolling ComfyUI back was indeed the practical workaround.

    This is unrelated to strict_compatibility. That setting only controls the model/checkpoint compatibility gate.

    Sorry that my previous reply made it sound as though a compatible Continuum release was already available. The compatibility adaptation has now been implemented and tested in the current development build. Until that build is released, rollback remains a valid temporary workaround.

    Thank you for identifying and clearly explaining the actual cause.

    dft78750707Aug 16, 2026· 4 reactions
    CivitAI

    Can we get a systemprompt for local llm like qwen3.5? I really have problems to get a working prompt structure.

    ukr8b3g201
    Author
    Aug 16, 2026

    Yes. MiniMax provides an official H3 prompt-writing skill here:

    https://github.com/MiniMax-AI/MiniMax-H3/tree/main/skills/h3-prompt-writing

    It supports T2VA, I2VA, FL2VA, L2VA, and Ref2VA, and includes the required prompt structures and reference-label rules.

    If your local Qwen setup supports agent skills, install or point it to the complete h3-prompt-writing folder.

    Otherwise, provide these files to the model as persistent system context:

    - SKILL.md

    - references/base-en.txt

    - references/ref-en.txt

    The whole folder should be used, not only SKILL.md, because the detailed formats and examples are stored in the reference files.

    For H3 Continuum, use the standard H3 structure generated by this skill, but divide the visual timeline into one section per Continuum chunk. In the currently tested workflow, that means one 5-second section per chunk.

    dft78750707Aug 16, 2026

    @ukr8b3g201 Thx, I know them and i have some good system-prompt already. but they do not work with your node, I tried to modify them, but either I get errors about reference images, or about the prompt structure and how every chunk needs to start. I got one running, but that creates 3 times the first video. What I was asking for was a system-prompt that puts out the exact structure that your node needs.

    Complement_CascadeAug 18, 2026

    @dft78750707 I found using https://civitai.red/models/2834106/minimaxh3-auto-prompter-v61?modelVersionId=3236905 and just adding break it into 3 shots of 5 seconds in the prompt worked for me

    ukr8b3g201
    Author
    Aug 18, 2026

    @Complement_Cascade Thanks, I checked it. MinimaxH3 Auto Prompter V6.1 is a useful optional front end for generating H3 prompts.

    For Continuum, I recommend changing “three shots of 5 seconds” to “three consecutive 5-second timeline sections,” because the word “shots” may encourage separate scenes or cuts.

    A safer instruction is:

    “Return exactly three consecutive timeline sections headed [0-5s], [5-10s], and [10-15s]. Keep them as one continuous sequence. Do not restart or repeat the subject, scene, or action at each section.”

    The generated result should still be reviewed before running, because this Auto Prompter is based on the standard H3 guide and is not specifically tied to the Continuum parser.

    Complement_CascadeAug 18, 2026

    @ukr8b3g201 nice change (I was previously having to RNG the prompts) but this change improved the outputs. and yes I normally use the autogenerated ones as a base and then change as appropriate

    ukr8b3g201
    Author
    Aug 18, 2026

    @Complement_Cascade Glad it helped. Using the generated prompt as a starting point and then refining the timeline sections manually is currently the safest workflow for Continuum.

    The important part is to keep the exact 5-second section structure and describe each section as a continuation of the same sequence, rather than as an independent shot.

    sdktertiaire2Aug 16, 2026· 3 reactions
    CivitAI

    hello, great workflow and work. thanks

    ukr8b3g201
    Author
    Aug 16, 2026

    Hello, thank you! I’m glad you found the workflow useful.

    SomeRandomUser23Aug 16, 2026· 3 reactions
    CivitAI

    Is there a proper way to do dialogue with ref2v? Everything works great until I try adding it. It will generate random speech outside the actual prompted dialogue

    ukr8b3g201
    Author
    Aug 16, 2026

    Use a stable speaker ID and place the exact dialogue inside an H3 dialogue tag in the timeline:

    [0-5s]

    <Subject 1> (S1) looks toward the camera and says in a calm young female voice:

    <d>[English] The last train should be here soon.</d>

    She finishes the sentence, closes her mouth, and remains silent. No other voices or background speech are heard.

    Keep these points in mind:

    - Put spoken words only inside <d>[Language] ...</d>.

    - Assign the speaker once, such as <Subject 1> (S1), and reuse (S1) later.

    - Keep each line short enough to finish inside its 5-second chunk.

    - Do not repeat or paraphrase the dialogue in the summary or soundscape.

    - Explicitly state that no other dialogue or background voices are present.

    - Avoid splitting one sentence across two Continuum chunks.

    Reference Audio can guide voice character, rhythm, or performance, but it does not guarantee an exact waveform or perfectly repeat the original speech.

    MiniMax H3 generates the audio rather than playing a fixed dialogue track, so exact scripted wording is not guaranteed and it may occasionally improvise extra words. For production work requiring exact dialogue, use external TTS and a separate lip-sync or audio replacement stage.

    CannBoyoAug 16, 2026

    I had the same issue until I realized I have to tell the model to shut the fuck up by adding something along the lines of "she finishes speaking" or as OP said "she finishes the sentence"
    otherwise there will be random bleeding of speech, works even if you sloppily use the " " for speech, but rather follow the <d>[Language] ...</d>

    SkuuurtAug 16, 2026· 2 reactions
    CivitAI

    I'm waiting for a "real" way to get a "real" reference video input (and not just the timeline video input) for this node.

    Having no reference video input in that sampler makes it impossible to get an identity (visual appearance + voice) in one shot. Unfortunately this is a deal breaker for me. Nonetheless, I appreciate the work you've done so far for the community, congrats !

    ukr8b3g201
    Author
    Aug 17, 2026

    Thank you. That is a valid distinction and a real limitation of the current node.

    Timeline Video is implemented as a chunk-local motion and composition guide. It is not equivalent to MiniMax H3’s native Reference Video input, which can provide visual appearance and associated audio/voice from the same source clip.

    A proper Reference Video implementation would need to preserve the native H3 video-reference semantics across all Continuum chunks, rather than slicing it like Timeline Video.

    I agree that this is important for identity and voice consistency. I will investigate adding a separate persistent Reference Video input, but I cannot promise a release date until its compatibility, memory usage, and interaction with Continuation have been verified.

    Thank you for clearly identifying the use case and for the kind feedback.

    VI6_D_DARK_KINGAug 17, 2026· 2 reactions
    CivitAI

    Is there a way to continue an existing run? Say I made a run with 3 Chunks closed Comfy and now I want to continue from Chunk 3, maybe with different Loras and different reference Images and audio and add more while keeping the exiting first 2 Chunks unaltered.

    Every time I'm tried it game gave a new run.

    ukr8b3g201
    Author
    Aug 17, 2026· 1 reaction

    Yes, an interrupted run can be resumed when Run Storage is enabled and the generation contract remains unchanged.

    Reopen the same workflow and keep the same Run Name, model, LoRAs, reference images/audio, resolution, seeds, prompts, and other sampling settings. Completed chunks should then be reused and only missing chunks generated.

    However, changing the LoRAs, reference images, or reference audio creates a new revision by design. Those inputs are persistent across all Continuum chunks, so the current version cannot keep Chunks 1–2 from the old contract while applying different references or LoRAs starting from Chunk 3.

    What you are describing is a “fork from Chunk N” feature: preserve the earlier chunks, inherit their continuation context, and start a new branch with a different conditioning contract. That is not currently implemented, but it is a valid feature request.

    Closing ComfyUI itself should not prevent resume, provided Run Storage was enabled and the same stored run identity can be found.

    VI6_D_DARK_KINGAug 17, 2026· 2 reactions

    @ukr8b3g201 OK. I know what the problem was.

    I had the seed set to Random. Not going to make that mistake again.

    And consider it a feature request then.

    On a side note. This really should be integrated into a Director node to reach its full potential.

    ukr8b3g201
    Author
    Aug 17, 2026· 1 reaction

    @VI6_D_DARK_KING Glad you found the cause. Yes, Auto Resume requires the seed and the rest of the generation contract to remain unchanged.

    I’ll consider “Fork from Chunk N” as a feature request: preserve completed chunks, inherit their continuation context, and continue under a new LoRA/reference/prompt contract.

    Regarding the Director node, could you clarify which specific node or project you mean? MiniMax has T2V/I2V-01-Director API models, but those are separate from local MiniMax H3 and H3 Continuum. If you mean a scene-planning or timeline orchestration node, that could be a useful layer above Continuum.

    VI6_D_DARK_KINGAug 17, 2026

    @ukr8b3g201 Yup. There are several, the one I use is the MiniMax H3 Director by seesee75. https://github.com/seesee75-commits/ComfyUI-MiniMaxH3-Director it includes an inbuilt prompt generator so I'm testing out 5s to 10s prompts at low res, before copying it to the H3 Continuum workflow. Oh and since I've got you. How hard would it be to implement the option of setting the length of each chunk?

    To avoid problems with the voice lines.

    ukr8b3g201
    Author
    Aug 17, 2026· 1 reaction

    @VI6_D_DARK_KING Thanks, that clarifies what you meant by a Director node. MiniMax H3 Director looks like a useful way to build and test the shot prompts before moving them into Continuum.

    H3 Continuum already has a chunk_seconds setting, but it currently applies the same duration to every chunk. Most of my runtime validation has been done with 5-second chunks, so I do not want to claim that longer chunk durations are equally validated yet.

    If you mean independently setting each chunk, for example 5s + 10s + 7s, that is not currently implemented. It affects more than the UI because the prompt ranges, H3 frame grid, continuation context, Run Storage contract, audio slicing, and final assembly all need to use exactly the same boundaries.

    Voice-line alignment is a valid reason for supporting it. I am currently reviewing the audio path and chunk timing for the next version, so I will treat variable per-chunk duration as a feature request, but I cannot promise an implementation or timeline yet.

    Thanks for sharing the Director workflow. Its compiled prompt output may also be useful for future interoperability with Continuum.

    byarloooarlooo279Aug 18, 2026· 2 reactions
    CivitAI

    Hi, i was trying to add a video in the "timeline video" node, but then it generated the exact same video as my video ref and ignored the image refs. I don't get what you mean by "timeline video". How different is it from the "reference video" in a standart Minimax H3 workflow? Is it possible to drive the motion of the video with a video reference, or "timeline video" means something else? Could you explain what I have to activate or swith in the Continuum node to use a video ref/timelin? Thanks for any info.

    ukr8b3g201
    Author
    Aug 18, 2026· 1 reaction

    Thanks for asking. Your result is consistent with the current Timeline Video design.

    Timeline Video is not a motion-only control, and it is not exactly the same as the standard MiniMax H3 Reference Video input.

    In the current Continuum implementation, the connected video is divided according to the Continuum timeline:

    - Chunk 1 receives the 0–5 second video slice

    - Chunk 2 receives the 5–10 second slice

    - Chunk 3 receives the 10–15 second slice

    - and so on

    Each slice is resized using Timeline Video Size, with Efficient — 0.4 MP as the recommended default, and is then supplied to H3 as a chunk-local video reference.

    However, an H3 video reference contains much more than motion. It can condition:

    - subject appearance

    - clothing

    - composition

    - background

    - lighting and colors

    - camera movement

    - body movement

    Therefore, Timeline Video can dominate the Picture references and produce something very close to the source video. Picture references remain connected across all chunks, but the current implementation cannot enforce “use the pictures only for identity and the video only for motion.”

    A standard H3 Reference Video supplies a reference-video block to a normal H3 generation. Continuum Timeline Video additionally divides a longer video into time-aligned slices and supplies the corresponding slice to each Continuum chunk.

    To activate the current feature:

    1. Connect the VIDEO output from Load Video to timeline_video.

    2. Set Timeline Video Size to Efficient — 0.4 MP.

    3. Keep the Picture references connected if required.

    4. Describe the pictures as identity references and the video as action/camera guidance in the prompt.

    There is no additional activation switch. Connecting timeline_video enables it automatically.

    If you specifically need character replacement while preserving only the source motion, the current Timeline Video implementation cannot guarantee that result. Reproducing too much of the original video is a known limitation rather than a missing switch.

    The video-conditioning direction is planned to be revised in a later version, after the current Driving Audio work. The planned design will distinguish source timing/motion guidance more clearly from identity replacement. The current Timeline Video should therefore still be considered an experimental video-reference mode rather than a dedicated motion-transfer system.

    byarloooarlooo279Aug 20, 2026· 2 reactions

    @ukr8b3g201 Thanks for the very detailed answer!

    Trying 3.4 now. And it's still very difficult for me to get rid of getting the ref video as final result (I tweaked during almost 1 day on 3.3 and starting tweaking on 3.4).

    Do you know any particular effective prompting? Or siwtching PROMPT FORMAT mode, does it get things better like swithing AUTO to FIXED? or something else to activate?

    ukr8b3g201
    Author
    Aug 20, 2026· 1 reaction

    @byarloooarlooo279 Thanks for testing 3.4.

    There isn't an extra switch that will make the Video Reference behave as motion-only guidance. Also, changing Prompt Format from Auto to Fixed will not reduce the Video Reference strength. If your prompt is normal continuous text, Auto already resolves it as Fixed internally.

    For V3.4, I recommend trying:

    Video Reference Size: Efficient — 0.4 MP

    Keep your Picture references connected

    Explicitly separate the roles in the prompt, for example:

    Use <Picture 1> as the primary reference for identity, face, hair, body appearance and clothing. Use <Video 1> only as guidance for motion, pose progression, timing and camera movement. Do not copy the person, clothing, background, lighting or colors from <Video 1>.

    You can also phrase the main action like:

    The person from <Picture 1> performs the motion shown in <Video 1>.

    However, this is still only prompting guidance. H3's Video Reference contains appearance, composition, background, lighting and motion information together, and Continuum currently has no separate Video Reference strength or true "motion only" control. So if the generated result still strongly reproduces the source video, that is a limitation of the current V3.4 reference-conditioning method rather than a missing setting.

    Your feedback is useful because it confirms that this can still happen in 3.4. I am looking at ways to separate identity conditioning from video motion/timing guidance more clearly in a future revision.