MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long. Conditioned on lyrics and a detailed music description, it generates structurally coherent songs with expressive vocals, evolving arrangements, and stable long-form audio quality.
MiniMax Music 3 combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE. The model produces 32 kHz, 16-bit stereo WAV audio.
Complete Songs with Long-Range Coherence
MiniMax Music 3 natively supports full-song generation up to five minutes. The model maintains musical themes, rhythm, vocal identity, and arrangement progression across long sequences, enabling complete structures such as intro, verse, pre-chorus, chorus, bridge, instrumental break, and outro.
Fine-Grained Music Control
The model accepts two complementary inputs:
Lyrics define the words to be sung and may include explicit section tags such as
[Intro],[Verse],[Pre-Chorus],[Chorus],[Post-Chorus],[Bridge],[Instrumental],[Solo], and[Outro].Music description defines the musical style, emotional progression, vocal performance, instrumentation, arrangement, and production profile.
For precise control, we recommend using a Structured Caption with three sections:
Global Metadata: genre, subgenre, BPM, key, scale, emotional progression, listening scenario, and production profile.
Vocal Details: vocal gender, timbre, performance style, harmony, backing vocals, and vocal effects.
Arrangement: primary and secondary instruments, section-level instrument evolution, groove, bass, percussion, textures, and spatial effects.
This representation allows the model to follow not only a global style, but also the musical development of the song over time.
