CivArchive
    Ace Step 1.5 XL Turbo and SFT - TEXT to AUDIO model with Ollama - v1.2
    NSFW
    Preview 120947173

    The Workflow was setup to have a clean "GUI" showing only parameters that matter, so you might want to toggle off Link visibility, like in above screenshot.


    V1.8 Ace Step 1.5. Turbo and SFT normal and XL model with Ollama. Text to Audio/Song (examples below).

    small update on Lora handling to better include Lora trigger words/phrases

    • Check this link with 1.5.Turbo XL Loras by Ryanontheinside: https://huggingface.co/ryanontheinside/models

    • 1.5 XL Loras can be loaded on any XL Slot (Turbo, SFT or merges), it improve the sound very well, like more authentic guitars for rock&metal, etc. Also song structures follow the genre better.

    • Loras require a trigger word that can be found on the main page per Lora of above link. Use the trigger word in the new node "Pretext" within the Input section of the workflow, added around 25 of those triggers to a Note in the WF for copy&paste.

    • See post in discussion below for more info.

    Key features:

    • Can use any Song, Artist as reference or any other description to generate tags and lyrics.

    • Will output up to 4 songs, each by Turbo, SFT, Turbo XL and SFT XL model (or any merge or Base model).

    • Keyscales, bpm and song duration can be randomized.

    • able to use dynamic prompts.

    • creates suitable songtitle and filenames with Ollama.

    • saves songs as MP3 with tags like lyrics, artist, genre, etc.


    V1.7 Ace Step 1.5. Turbo and SFT normal and XL model with Ollama. Text to Audio/Song.

    • replaced the "save Mp3" nodes to allow ID3 tags to include data like artist name, bpm, genre, lyrics, etc.

    • added a feature to save all relevant data (tags, lyrics, etc.) as a separate text (.txt) file

    • audio render processing remains unchanged to previous version


    V1.6 Ace Step 1.5. Turbo and SFT normal and XL model with Ollama.

    updated the settings for XL models and added a 3rd System Prompt for tags to chose, with more descriptive song descriptions.

    1.5 XL SFT pipeline now has an "Adaptive Projected Guidance" node and negative prompt.

    ** See below some tips which model and settings to start with.


    V1.5 Ace Step 1.5. Turbo and SFT normal and XL model with Ollama:

    • setup to create up to 4 tracks in a run, 2x Ace1.5 and 2x Ace1.5 XL, each with Turbo and SFT model, to compare (can be individually switched on/off)

    • VAE changed to tiled Audio VAE decode, uses less Vram.


    V1.2 Ace Step 1.5. Turbo and SFT model with Ollama Text to Audio/Song

    • small update to GUI, system prompts and SFT sampler "engine"


    V1.0 Ace Step 1.5. Turbo and SFT model with Ollama Text to Audio/Song

    Ace Step uses TAGS and LYRICS to create a song. These can be generated by Ollama or by own prompts.


    Download Files:

    Ollama Models, required for tags, lyrics and songtitle, you can choose 1,2 or 3 different models, tags and lyrics might need a bigger model >7b, songtitle can use a smaller model:


    Alternative Turbo Models and merges (normal, non XL) :


    GGUF Models "normal" and XL: https://huggingface.co/Serveurperso/ACE-Step-1.5-GGUF/tree/main


    Which models to start with ?

    • If you just want to try it first before downloading all those models, start with 1 model only, recommend the Turbo or Turbo XL model.

    • My current choice for normal model: Turbo-SFT merge_ta_0.5 & Turbo-Shift1, using these settings:

      • Turbo-SFT_merge model with sampler: er_sde, scheduler: beta57 (or beta), 22 steps

      • Turbo-Shift1 model with sampler euler, scheduler: normal, 138 steps

    • XL Model settings:

      • XL Turbo-SFT merge model: sampler: er_sde, scheduler: sgm_uniform, 40 steps

        • alternative: sampler: res_s2, scheduler beta57 (requires RES4LYF custom nodes)

      • XL SFT model: sampler: euler (or res_2s), scheduler: normal, 46 steps, CFG = 7.3, Adaptive Projected Guidance: eta = 1.05, norm_thresh= 1.3, momentum=0.0. Increase norm_thresh as the main parameter. These settings deliver "stabil" output for XL SFT,Base and their merges. The merges sound way better, pure SFT or Base introduce a lot of noise. I bypassed ModelSamplingAuraflow (see node next to model loader node). I think the base-turbo XL model merge fits well in that slot.

    • Disable "generate_audio_codes" in "TextEncodeAceStep" node to get different results, it works very well for many genres and reduces process time.

    • Ollama Model: Llama-3-NeuralDaredevil-8b-abliterated

    More infos on models see thread below in discussion.


    Save Location:

    • ๐Ÿ“‚ ComfyUI/

    • โ”œโ”€โ”€ ๐Ÿ“‚ models/

    • โ”‚ โ”œโ”€โ”€ ๐Ÿ“‚ diffusion_models/

    • โ”‚ โ”‚ โ””โ”€โ”€ acestep_v1.5_turbo.safetensors

    • โ”‚ โ”œโ”€โ”€ ๐Ÿ“‚ text_encoders/

    • โ”‚ โ”‚ โ”œโ”€โ”€ qwen_0.6b_ace15.safetensors

    • โ”‚ โ”‚ โ””โ”€โ”€ qwen_4b_ace15.safetensors (or 1.7b)

    • โ”‚ โ””โ”€โ”€ ๐Ÿ“‚ vae/

    • โ”‚ โ””โ”€โ”€ ace_1.5_vae.safetensors


    Custom Nodes used:

    optional (use Beta57 scheduler for a bit more punch, requires RES4LYF): https://github.com/ClownsharkBatwing/RES4LYF


    Examples various styles:

    With Lora:

    No Lora:


    Ollama help:

    1. Install Ollama from https://ollama.com/

    2. download a model: Go to a model page, chose a model , then hit the copy button, i.e. https://ollama.com/mirage335/Llama-3-NeuralDaredevil-8B-abliterated-virtuoso

    3. open terminal and paste the model name, i.e.: ollama run huihui_ai/qwen3-vl-abliterated

    4. model will be downloaded and can be selected in green comfy node "Ollama Connectivity". Hit "Reconnect" to refresh.

    Description

    Ace Step 1.5. Turbo and SFT model with Ollama for song-tags and lyrics

    Gui, System prompts and small engine update

    FAQ

    Comments (7)

    IafunFeb 14, 2026
    CivitAI

    Hi, I have a problem; when I try to load the workflow, I get this error. :TypeError: can't access property "readOnly", e.inputEl is undefined

    tremolo28
    Author
    Feb 14, 2026

    Hi try update All in comfyui and clear browser cache.

    galvo3d577Feb 20, 2026
    CivitAI

    Hi! Is it possible to have the wildcards also?

    tremolo28
    Author
    Feb 20, 2026ยท 1 reaction

    Hi, you can search CivitAt for wildcards. It is an own category, like here: https://civitai.com/search/models?modelType=Wildcards&sortBy=models_v9&query=wildcards

    Or create your own wildcards, it is just text (.txt) files containing prompts

    TribbleConApr 2, 2026
    CivitAI

    Are lyrics required? I'm looking to generate chiptunes and melodies.

    tremolo28
    Author
    Apr 2, 2026

    no, you can just say "instrumental".

    XIIAIApr 16, 2026ยท 1 reaction

    You can throw the normal tags for crafting the flow of the song, [Instrumental], [Intro], [Verse], [Pre-Chorus], [Chorus], [Refrain], [Bridge], [Outro], etc in the prompt to define your song's pattern. I usually fill the line below each stage with a couple underscores to indicate content to the the process so it doesn't just stack all the stages up in one go and then generate random after that. [guitar solo] and [instrument solo] also work. despite the limitations indicated in the official prompting documentation, the generation does handle abstract ideas in the instrumentation prompt before the lyric stage.

    Workflows
    ACE Audio

    Details

    Downloads
    753
    Platform
    CivitAI
    Platform Status
    Available
    Created
    2/13/2026
    Updated
    8/15/2026
    Deleted
    -

    Files