The complete bill for the song I generated last night was the electricity it drew. Maybe a few cents.
Full track. Vocals, drums, bass, a bridge that actually goes somewhere. I made it on the GPU I already own, with model weights anybody can download, and there was no account, no credits counter, no monthly plan ticking down. If I want forty more songs today, I can have forty more songs today.
That sentence wasn't true a year ago. Free AI music meant demos that sounded like a karaoke machine falling downstairs. But in the past few weeks the open-source side quietly crossed a line, and "just run it yourself" stopped being a hobbyist's flex and became the reasonable option.
If you've been paying for AI music, this is worth fifteen minutes of your attention. (And if you're a Suno subscriber, the caps landing September 3 — 20 downloaded songs a month on Pro, even for your existing library — probably already got your attention. That news is what sent me down this road. It's just not the interesting part. The interesting part is what's now free.)
The short version
Four open-source tools cover the whole pipeline, end to end, at zero monthly cost:
| Tool | What it does | What it needs | The catch |
|---|
MiniMax Music 3 | Full songs with vocals, up to 5 min | 16 GB VRAM comfortable, 8 GB claimed | Community license — read it before commercial use |
ACE-Step 1.5 | Covers, section edits, stem splits, style transfer | Under 4 GB VRAM | Quality trails the big closed models |
Foundation-1 | Loops and samples that follow your BPM, key, bar count | Modest GPU | Built for DAW workflows, not one-click songs |
MuScriptor | Turns any recording into per-instrument MIDI, vocals included | Small version fits on a phone | Weights are non-commercial |
Generate, arrange, transcribe, polish. All local, all free to run. Let me walk you through it, because each piece is worth knowing on its own.
Start here: MiniMax Music 3
This is the one that changed the math. MiniMax released Music 3 with downloadable weights, and it does the whole job: you describe the sound, paste in lyrics, and it hands back a finished song up to five minutes long. Clean 32 kHz stereo. It shipped with day-one support in ComfyUI, the tool most people already use to run open image and video models on their own hardware.
What I like is how you talk to it. The description field works in layers: the global stuff first (genre, mood, tempo, key), then the voice (male or female, smooth or rough, where it sits in the mix), then the arrangement — which instruments carry the intro, where the bridge drops out. Lyrics take the usual [verse] and [chorus] tags, backing vocals in brackets, even breaths and pauses if you want them. Fix the seed and the song reproduces exactly. Change the seed, get a different take of the same idea.
It is not fast. A one-minute song renders in three to four minutes on a 16 GB card. The model comes in three sizes — roughly 10 GB, 5 GB, and 2.5 GB — so smaller cards aren't locked out; MiniMax claims 8 GB works if you let it stream memory. But nobody's charging you by the minute, so who cares.
And the honest part, because you deserve it: it is not Suno V5. Early hands-on comparisons put it around the polish of Suno's older mid-generation models. Clean, genuinely pleasant, occasionally great, a step behind the best closed systems on vocal nuance and instrumental variety. It is the strongest open music model that exists today. Both sentences are true at once.
Free used to mean "worse." Now it just means "yours."
The rest of the kit
MiniMax Music 3 does exactly one thing: text to finished song. No covers, no editing an existing track, no "make it sound like this." That's what the other three are for.
ACE-Step 1.5 is the Swiss army knife. Generate a song in a referenced style, repaint just the chorus, lift individual stems out of a mix, strip vocals for an instrumental. You can fine-tune it on a handful of songs to build a consistent house sound — genuinely useful if you want an audio identity rather than whatever the prompt gods hand you that day. It runs in under 4 GB of VRAM and renders a full song in under ten seconds on a two-generation-old GPU. For my money, this is the most commercially useful item on the list.
Foundation-1 comes at it from the producer's side. Instead of finished songs it makes loops and phrases that obey the BPM, key, and bar count you specify — the actual building blocks of a DAW workflow. Stack a bassline, a pad, a drum groove, export to MIDI, keep editing. If you have anyone on the team who can assemble audio, this fits straight into what they already do.
MuScriptor runs the pipeline backwards. Feed it any recording and it transcribes what every instrument played into MIDI — including the vocal line. It comes out of Kyutai and Mirelo, serious audio-research shops, and the small version is 412 MB; they say it fits on a phone. The practical loop: generate a track, pull the MIDI, fix the flat parts by hand, replay them with better sounds. That last-mile polish is exactly where generated music usually loses people, and this closes it.
The catches
I'd be selling you something if I stopped at the highlight reel.
- The quality gap is real. For a hero brand asset, Suno's latest is still better. Free closes the gap on volume and ownership, not on the very top end.
- The downloads are chunky. Realistically you'll pull 12–20 GB of model files before the first note plays. One-time cost, but plan the disk.
- "Free" is a spectrum. MiniMax ships a community license with conditions worth reading before commercial use. MuScriptor's weights are explicitly non-commercial. ACE-Step asks for care around protected styles. Check each one against how you'll actually use the output.
- Setup is an afternoon, not a click. Install and configure territory. For a solo creator that's a weekend project; for a team it's a one-time setup cost that pays back the first month nobody thinks about export quotas.
Set against $8–24 a month forever, plus — as of September — a hard ceiling on how many songs you're allowed to keep, those catches start looking like rounding errors.
So should you run your own?
Loading diagram…
My read: if music is a garnish — a handful of tracks a year — stay where you are. The moment audio becomes infrastructure (weekly videos, a podcast network, ads across channels, client deliverables), generating it yourself on hardware you own is now the sensible default, and the paid clouds are the workaround.
There's a bigger pattern here too, and it's not about music. Image models, voice models, video models — the same fork keeps appearing, and each time the free local option arrives closer behind the paid one. The people learning to operate AI models as owned infrastructure, the way they once learned to run their own servers, are building a cost structure their competitors can't match.
If that's a fork you're staring at — music or otherwise — book a call and I'll help you figure out which side of it you should be on.