ACE-Step 1.5 just dropped in ComfyUI. Full songs in under 10 seconds. Less than 4GB VRAM. Open-source music generation just got serious.

Feb 3, 2026 · 5:34 PM UTC

30
62
524
75,409
The speed: ~1 sec on RTX 5090, <10 sec on RTX 3090. The quality: scoring beyond most commercial models on coherence benchmarks. 50+ languages. Runs locally on your hardware.
1
1
32
2,689
How it works: an LM plans the full song structure—lyrics, sections, arrangement—then a Diffusion Transformer generates the audio. Chain-of-thought reasoning keeps long-form compositions coherent.
2
16
3,533
Sort replies: Relevant Recent Liked
Replying to @ComfyUI
Bye, bye SUNO, bye bye SPOTIFY.. I've tried it and I can make 10 amazing songs in 20 minutes.
3
1,714
Replying to @ComfyUI
Tried it already! So cooool.
1
473
Replying to @ComfyUI

ALT 2sec Spongebob GIF

1
318
Replying to @ComfyUI
Quality: Heartmula (1x 5090 GPU, 48+ VRAM, 2-3 mins gen took ~5 mins) Form factor (VRAM, speed): ACE-Step 1.5
1
399
Replying to @ComfyUI
It's now possible to generate so much material that double the playback speed is becoming increasingly important :) ACE-Step 1.5
4
563
Replying to @ComfyUI
Not really. It does not adhere to lyrics for @#$%.
4
407
Replying to @ComfyUI
lowering the barrier to 4gb vram is a huge win for local production. moving away from cloud api costs and latency means independent builders can actually scale these workflows now. great to see this landing in comfyui.
3
545
Replying to @ComfyUI
@ComfyUI just updated comfyui and it won't start. My friend has the exact same problem.
2
2
1,051
Replying to @ComfyUI
Last week, I almost hit the buy button for home music generation Place funds in jar for 5090
3
672
Replying to @ComfyUI
Awesome! Let’s make some Comfy music!
1
552
Replying to @ComfyUI
Does it do "covers" of uploaded audio like Suno does? If not I'm not interested.
2
435
Replying to @ComfyUI
wait less than 4 GB vram???
2
398
Replying to @ComfyUI
Oh nice. I am gonna have to look at that. I was just talking to a friend about trying some music generations using ComfyUI.
1
555
Replying to @ComfyUI
Really impressed with online demo. Thanks for the release.
1
367
Replying to @ComfyUI
I need to see some voice notes samples to music using this 🤩
400
Replying to @ComfyUI
is this better quality or just for lower VRAMS?
1
373
Replying to @ComfyUI
x.com/nelvOfficial/status/20… not sure if the way it works is similar to suno or not, but its mindblowing to know that musicAI actually generates an image first before the sound
"Katsura Kurumi Explains - (Tech Edition)" "S2-EP01: Suno — The Probability Orchestra" 1/ #AI #KatsuraKurumi
762
Replying to @ComfyUI
It's great, i need a Workflow to repaint
1
92
Replying to @ComfyUI
The <4GB VRAM requirement is impressive for full song generation. ACE-Step uses latent diffusion to minimize memory footprint while maintaining coherent structure over longer sequences. Perfect for local workflows without cloud API dependencies.
35
Replying to @ComfyUI
🔥🎶
418
Replying to @ComfyUI
I'm confused, it's not there?
1
356
Replying to @ComfyUI
can you clone you own voice with this one? @grok
1
74
Replying to @ComfyUI
1 hour to get a 4 minutes song on my 5060
104
Replying to @ComfyUI
Why this sounds so clear and outstanding? It's amazing
327