Audio & Voice

HeartMuLa

The Stable Diffusion moment for music

Open ✓Model

What it is

A family of open music foundation models: a multilingual lyrics-and-tags-conditioned song generator (3B open, 7B internal), the HeartCodec 12.5Hz music codec, a lyrics transcriber, and a CLAP-style audio-text aligner. Generates full songs with vocals at roughly real-time speed.

Why it's interesting

Music generation has been dominated by closed products like Suno and Udio. A genuinely Apache-2.0 release — code and weights, commercial use included, relicensed explicitly in January 2026 — is the moment the open music ecosystem has been waiting for.

Use cases

  • Full-song generation from lyrics and style tags
  • Music-AI research and fine-tuning on custom catalogs
  • Lyrics transcription

Who it's for

Music-tech developers, generative-audio researchers, indie creators

Setup

Moderate. A CUDA GPU (the 3B model fits prosumer cards), PyTorch, and tens of GB for weights from HF or ModelScope

Limitations & cautions

Not yet optimized for streaming, no reference-audio conditioning, and the 7B 'Suno-comparable' variant remains unreleased. Industry-wide training-data copyright questions are unresolved here too.

Editorial takeaway

Watch what indie musicians do with this in a year. Open weights change who gets to experiment.

Related & alternatives