MiniMax Music 3.0

Audio·minimax-music·by Minimax

MiniMax Music 3.0 creates studio-quality songs from lyrics and style prompts, with support for instrumental tracks, automatic lyric generation, and MP3 output.

Runtime (p50)
1m
Estimated price
$0.15
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "minimax-music-03",
    "version": "0.0.1",
    "input": {
        "lyrics": "[Verse 1]\nOne clean call and the engine takes flight,\neach::labs turns the idea into light.\nImage or voice, a scene in motion,\nEvery model runs on the same devotion.\n\n[Chorus]\neach::labs, the current, the spark, the run,\nFrom a single prompt to a product done.\nBuild it, ship it, watch it grow,\nEverything you imagined ready to go.\n\n[Verse 2]\nNo keys to juggle, no stack to fight,\nDocs that make sense at two in the night.\nBuilders and dreamers moving as one,\nWe are in it with you until it is done.",
        "prompt": "Anthemic electronic pop, layered male and female vocals, driving beat, bright synths, punchy bass, uplifting build into a wide chorus, confident and warm, 120 BPM.",
        "output_format": "url",
        "audio_settings": {
            "format": "mp3",
            "bitrate": 256000,
            "sample_rate": 44100
        },
        "is_instrumental": false
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    MiniMax Music 3.0 Overview

    MiniMax Music 3.0 is a next-generation music-generation model from Minimax that turns creative concepts and optional lyrics into fully produced, studio-quality songs in a single pass. As the flagship of the minimax-music family, it replaces earlier versions such as MiniMax Music 2.6 and is now the recommended text-to-music model for both API users and open-weight workflows. Its primary differentiator is long-form, structurally coherent songs of up to about five minutes, with expressive vocals, evolving arrangements, and stable audio quality over the full track. Integrated into each::labs, MiniMax Music 3.0 lets you generate vocal songs, instrumentals, or automatically written lyrics and export them as standard audio formats suitable for streaming, editing, and distribution.

  • Capabilities

    Capabilities

    • Generates complete songs up to around five minutes from text prompts and optional lyrics, including composition, arrangement, performance, and production in one pass.
    • Supports vocal songs with structured, user-supplied lyrics using section tags such as [Verse] and [Chorus], enabling precise control over narrative and hook placement.
    • Creates pure instrumental tracks with no vocals from descriptive style prompts, suitable for background music and commercial soundtracks.
    • Offers automatic lyric generation mode, where MiniMax Music 3.0 writes lyrics to match the prompt’s concept and then performs them as a finished song.
    • Delivers commercial-grade audio quality with clearer mix, more realistic instrument articulation (including slides and legato), and improved vocal expressiveness and pronunciation.
    • Maintains long-form structural coherence, with evolving arrangements and stable audio quality across several minutes of music.
    • Provides an officially documented MiniMax Music 3.0 API with model IDs music-3.0 and music-3.0-free, rate limits, and pay-per-song pricing.
    • Ships as open weights for local and custom deployments, enabling advanced users to integrate the model into their own pipelines and tools.
  • Use cases

    Use Cases for MiniMax Music 3.0

    For independent creators and artists, MiniMax Music 3.0 can turn a fully written lyric sheet into a polished demo with expressive vocals and a complete arrangement, simply by tagging sections and describing the desired style. A sample workflow might use: “Pop-rock band track with live drums and guitars, [Verse] storytelling about leaving home, [Chorus] big anthem hook around ‘run towards the light’.”

    Marketers and brands can use MiniMax Music 3.0 to generate instrumental beds that evolve over several minutes, matching campaign mood and pacing. An example prompt: “Instrumental only, modern electronic chill with soft pads and light percussion for a 5-minute product launch livestream intro.”

    Developers integrating music into apps or games can call the MiniMax Music 3.0 API to dynamically create background tracks or theme songs based on in-app events. For instance: “Dark orchestral score with subtle choir for a boss fight scene, structured intro, build, and finale over 3 minutes.”

    Designers and video editors can rely on long-form outputs to avoid looping short clips, using prompts such as: “Cinematic ambient track with gentle piano and evolving textures, timed to a 4-minute travel montage.”

  • Tips & tricks

    Tips and Tricks

    MiniMax Music 3.0 responds best to producer-style prompts that describe genre, mood, instrumentation, and structure, combined with clearly tagged lyrics. Use section labels like [Intro], [Verse], [Pre-Chorus], and [Chorus] on their own lines to guide song form and help the model maintain coherent structure over several minutes. When using the MiniMax Music 3.0 API, separate your music description (style, tempo, arrangement) from the lyrics, and specify whether you want a vocal track or instrumental-only output. For automatic lyric generation, start with a concise creative concept and emotional direction, then let the model write and perform the song.

    Example prompts:

    “An upbeat K-pop track with bright synths, tight drums, and energetic female group vocals; [Verse] lyrics about chasing dreams, [Chorus] a catchy hook built around the phrase ‘lights up the sky’.”

    “A melancholic piano-led pop ballad with intimate male vocals and subtle strings; auto-generate lyrics about a relationship slowly fading over time.”

    “Instrumental only: cinematic indie-folk bed with warm acoustic guitars, soft percussion, and evolving textures for a 4-minute brand video soundtrack.”

  • Technical spec

    Technical Specifications

    • Max song length: up to about five minutes (roughly 300 seconds) per generation.
    • Audio output: 32 kHz, 16-bit stereo WAV in the reference implementation; MP3 and URL outputs are commonly supported by hosting platforms.
    • Input formats: text prompts plus lyrics, with structure tags like [Verse] and [Chorus] placed on separate lines.
    • Architecture: hierarchical autoregressive system with an 8B global language model for long-range musical structure, a 0.6B local model for acoustic detail, and Flow-Matching / Flow-VAE synthesis for audio.
    • Processing time: community reports indicate several minutes of compute for multi-minute songs on typical GPUs when running locally; hosted APIs return finished tracks asynchronously.
  • Things to be aware of

    Things to Be Aware Of

    Because MiniMax Music 3.0 is optimized for full songs, ultra-short stingers or tightly timed sound logos may require more trial and error to align exact hit points. Lyric-heavy projects depend on clear structure tags and clean text; ambiguous or untagged lyrics can lead to less predictable phrasing or section lengths. While the MiniMax Music 3.0 API documents rate limits and flat per-song pricing, large-volume workloads still need careful budgeting and batching strategies. When running the open weights locally, hardware requirements are non-trivial, and community feedback suggests multi-minute songs can take several minutes of GPU time per generation on consumer cards. As with any AI music-generation system, users should review licensing and usage terms before commercial release.

  • Key considerations

    Key Considerations

    MiniMax Music 3.0 is designed for users who want full songs rather than short stingers or loops, making its five-minute capacity especially valuable for creators and marketers working on complete tracks. Accessing the MiniMax Music 3.0 API requires an account, API key, and adherence to documented rate limits (for example, 120 requests per minute on the paid tier and 3 RPM on the free tier). Compared with older Minimax music-generation models, MiniMax Music 3.0 improves semantic understanding of prompts, instrument articulation, and vocal control, which makes it a better choice when lyric fidelity, arrangement variety, and “commercial-grade” audio quality matter. For very short cues, sound effects, or highly edited stems, more specialized audio tools may still be preferable.

  • Limitations

    Limitations

    MiniMax Music 3.0 is not a full-featured DAW or stem editor; it focuses on generating complete mixed songs and instrumentals rather than exposing multitrack stems or MIDI exports. The model’s strengths lie in text-guided composition and performance, so extremely fine-grained control over note-by-note detail or post-hoc arrangement edits is limited compared with dedicated production tools. Reference-audio cover fields are explicitly documented as belonging to a separate model, meaning MiniMax Music 3.0 API calls should not rely on cover-specific parameters. Finally, while the model achieves commercial-grade quality for many genres, niche styles and highly experimental sound design may require additional processing or layering outside the MiniMax Music 3.0 API.

Related models

4 models