MiniMax Music 3: State of the Art Open Weight Music Generation
Blog post from Comfy
MiniMax Music 3 is an open-weight music generation model available in ComfyUI that produces complete songs of up to five minutes from lyrics and a music description, delivering 32 kHz, 16-bit stereo audio. Designed to maintain long-range musical coherence, it can preserve melody, rhythm, vocal identity, and evolving arrangements across structures such as verses, choruses, bridges, instrumental sections, and outros. Its hybrid architecture combines an 8B global language model for overall song structure, a 0.6B local model for acoustic detail, and continuous Flow Matching and Flow-VAE synthesis components intended to improve vocal articulation, instrumental texture, and temporal continuity. Users can control outputs through tagged lyric sections and Structured Captions specifying metadata such as genre, BPM, key, emotional progression, vocal characteristics, instruments, and section-level arrangement changes. MiniMax also provides a caption-rewriter skill to expand short prompts into structured descriptions, while the model and workflow can be accessed through ComfyUI version 0.33.0 or later, Comfy Cloud, and the published model weights repository.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 3 | 5,068 | 1,020 | 229 | -34% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.