ACE-Step 1.5 is an open-source music model similar to Suno, Udio, Mureka, Lyria, etc.
Workflows are embedded in the videos
Just drag and drop them into ComfyUI to import it (it includes an optional group of nodes to load a static image and turn it into a video along with the generated song).
Features
Lightweight: The model runs locally with less than 4GB of VRAM.
Fast: a 2 minutes song can be generated in a minute or so in a low-end GPU (<= 8GB VRAM).
Uncensored: you can prompt any lyrics you want.
Multiple languages: over 50 languages are supported officially.
Models
SFT 1.7B AIO checkpoint
I've merged the following models into a single checkpoint file:
acestep-v15-sft.safetensors: the SFT model, which sounds better than the Turbo version, while requiring only a few more steps.
qwen_0.6b_ace15.safetensors: mandatory text encoder.
qwen_1.7b_ace15.safetensors: additional 1.7B text encoder, which is many times faster than the 4B alternative while being almost as good in my own experiments).
ace_1.5_vae.safetensors: the VAE.
Simply download it to your ComfyUI/models/checkpoints folder.
Tips
Prompts
The model expects 2 optional prompts: caption/tags and lyrics/structure.
Caption/tags
Guide the model to what kind of genre and instruments you want in your music, as well vocals, mood, era, mixing style, etc.
Genres:
jazz,rap,rock,metal,hip hop,bossa nova,electronic,synthwave,blues,reggae, etc. A potential list of all 178k genres used during training can be found here.Instruments:
acoustic guitar,piano,drums,bass,synths,electric guitar,violin, etcVocals:
raspy male vocal,young female vocal,duet harmonies,whispering child, etcSee more tags in the official tutorial.
Either comma-separated tags, or natural language.
Multiple genres may be provided, but conflicts will likely harm the quality.
Even small changes will impact the result substantially.
Lyrics/structure
The model will sing better if each line contains between 6 and 10 syllables.
It's recommended to provide structure tags to organise your lyrics. Examples:
[Intro],[Verse],[Chorus],[Bridge],[Instrumental],[End], etc.
Add an empty line between structure blocks.
Inside a structure tag, you may add other hints. Examples:
[Intro - Dreamy],[Chorus - Layered vocals],[Instrumental - Guitar solo],[Bridge - Whispering], etc.
Some singing techniques/effects are recognised. Examples:
ACE-Step is heeere: hold the note for longer.For your mind (your mind): backing vocals.Stand up and SHOUT: sing with more power.[pt] Obrigado, amigo: switches to a different language[whispering] Don't be afraid: attempt to add said effect.
Metadata
All metadata are an effort to guide music attributes, but they might be overridden by the prompts.
The most relevant are listed below:
bpm: beats per minute, determines the tempo. Common distribution: slow songs 60–80, mid-tempo 90–120, fast songs 130–190.duration: target duration in seconds. The model officially supports 10s-600s. Short songs (30–60s) and medium length (2–4min) are stable; very long generation may have repetition or structure issues.timesignature: 4/4 (most common), 3/4 (waltz), 6/8 (swing feel).language: choose one of the many supported languages for the lyrics.keyscale: Affects overall pitch and emotional color. Usually Minor = darker mood; Major = brighter mood.generate_audio_codes: when enabled (recommended), spends much more time on the text encoder conditioning to improve song quality substantially.
Sampling
Most of the music (melody, harmony, cadence, etc) comes from the conditioning, so tweak sampling parameters to explore variations.
steps: the SFT model requires at least 20, but I recommend 30-50 for good results. Sometimes requires 50-100+ to improve failed parts.cfg: 1.0 is good enough, even for the SFT model. Increasing it to 2.0 seems to improve vocals while reducing presence of instruments. Over 2.0 starts to harm the output (do your own experiments though).sampler: my favourites:sa_solver_pece,heun,dpmpp_sde,uni_pc_bh2,euler, etc.scheduler: my favourites:beta,simple,kl_optimal, etc.
Description
ScragVAE — Improved VAE Decoder for ACE-Step 1.5
This is a drop-in replacement for ACE-Step 1.5 standard VAE.
Full credits to the original author: https://huggingface.co/scragnog/Ace-Step-1.5-ScragVAE
FAQ
Comments (15)
Hi. I failed creating a workflow for a music cover workflow for this model. have you ever tried? if so, Can you please share it with us?
Edit: What I actually mean is a workflow that can handle Music Audio covers/remixes/remasters with an initial audio file input.
literally 4 nodes. Go to workflows in COMFUI and use official workflow
@gambikules858 thx, i guess. already did that, vut the reaulta were horrible, even after changing and playing with many variables. do could you maybe Post a working example with this model?
@Adaptalab0r The workflow is embedded in my images in the showcase/gallery. I also attached it in the "Optional files" section underneath the Download button. Let me know if that's what you were looking for.
@SimplesmenteIA thx. i'll give it a shot and report back :)
@SimplesmenteIA Sorry for the misunderstanding. I don't mean cover art. What I actually mean is a workflow that can handle Music Audio covers/remixes/remasters with an initial audio file input.
@Adaptalab0r Got it. These features are indeed supported by the model, but you would need to install custom nodes in ComfyUI. Try this one: https://github.com/ryanontheinside/ComfyUI_RyanOnTheInside
@SimplesmenteIA thx looks promising. i'll try
Thanks
how many steps for the non-turebo version? I can't seem to get clear audio. it is usually distorted.
You can find the parameters in the embedded workflows (I also attached it as one of the downloadable files) - but I mostly use at least 30-50 steps. It might sound too many, but that's not much slower than the Turbo version.
Passando para agradecer. Outra coisa: mano, coloca um nome que defina melhor o modelo. Parece um modelo genérico e não é. É o modelo que você juntou e colocou aqui. Tem que ter a tua assinatura. Assim fica muito fácil de achar. Também queria uma versão quantizada do modelo em GGUF. Seria top.
Opa, valeu, mas o modelo não é meu, apenas juntei ele num checkpoint e subi aqui no Civitai, por isso usei o nome original. Já se fosse um finetune ou merge, aí sim eu teria renomeado. Se eu esbarrar numa versão em GGUF, eu compartilho. Abraço!
Thanks for this! The Turbo AIO was soooo... unpredictable. Getting much better results with your checkpoint.
I am deeply impressed!
The outputs are mostly what I had in mind, wayyyy more then MiniMax Music.
With the prompting I sometimes don't know, how to phrase the things. Will "chest-voice belt" be understood? "high octane snares"? "wailing guitars"?
A four line Chorus all in CAPS often loses power in the third line. Is there a trick?
I would like my Roleplay-Charakter to sound like Emily Armstrong. All the chorus.
Or just asking: Which genre is "Heavy is the crown" from Linkin Park? Simply Heavy Metal?
It's worth spending time and learning all this!