The entry threshold has been lowered to a minimum: the junior model requires less than 6 GB of video memory, and, depending on the think mode settings, generation can take from 2 to 10 seconds - this is already the level of commercial solutions.
🟢 What they did?
The developers have assembled a hybrid of a language model that turns a prompt into a composition sketch: it outlines the structure, comes up with lyrics and metadata, and DiT, which is responsible for the sound. The logical core of this entire system is based on Qwen3.
ACE-Step v1.5 can generate tracks from 10 seconds to 10 minutes long, with up to 8 tracks at a time. There are more than 1000 instruments in the database, and the system understands lyrics in 50 languages.
The authors have prepared a whole set of models for different amounts of VRAM:
➜ Less than 6 GB: without the LM module, only the sound engine works.
➜ 6-12 GB: a lightweight version of LM (0.6B).
➜ 16 GB and above: a full-fledged model with 4 billion parameters, which best understands the context and delivers maximum quality
.
When launched, ACE-Step v1.5 automatically selects a model and parameters suitable for the hardware. Detailed information on configurations can be found here.
ACE-Step can do much more than just turn text into a melody. You can give it an audio example to copy the style, make covers, correct parts of already finished tracks, or generate an accompaniment for vocals.
The most interesting feature is the ability to create LoRA. To feed the model with your own style, just 8 tracks are enough. On the 30th series RTX with 12 GB of memory, this process will take about an hour.
Everything is in order with the deployment, the developers have prepared a portable build, and for ComfyUI they have already written all the necessary nodes and workflows.
Project page, Model, Paper, Demo, Discord community, GitHub • #AI #ML #Text2Music #AceStudio #StepFun
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore