Lyria 3.5

Audio Gen·gemini·by Google

Lyria 3.5 generates full-length songs from text or image prompts, returning high-fidelity MP3 audio with vocals, lyrics, and song structure.

Runtime (p50)
1m
Estimated price
$0.08 per execution
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "lyria-3-5",
    "input": {
        "prompt": "A cinematic pop song with orchestral strings, warm analog synths, fingerpicked electric guitar, deep bass, and layered percussion. Expressive vocals shift between a soft vulnerable verse and a soaring powerful chorus with rich harmonies. The arrangement builds from a minimal piano intro through layered textured sections into a full orchestral climax, then strips back for an intimate outro. Polished high-fidelity stereo production, emotionally dynamic, golden hour atmosphere. 110 BPM."
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Lyria 3.5 Overview

    Lyria 3.5 is Google DeepMind’s latest Google music-generation model in the Gemini family, designed to turn natural-language or image prompts into full-length, studio-style songs with vocals, lyrics, and structured arrangements. It is exposed to developers through the Lyria 3.5 API as part of the Gemini API model lineup, and to creators through Gemini surfaces and Flow Music-style experiences. Within a couple of minutes, Lyria 3.5 can generate coherent tracks with verses, choruses, bridges, and synchronized lyrics, returned as high-fidelity stereo MP3 (and optionally WAV in some surfaces). A key differentiator is its focus on realistic, emotionally nuanced vocals and structural coherence, rather than short clips, making Lyria 3.5 particularly well suited for song demos, content soundtracks, and rapid prototyping from text.

  • Capabilities

    Capabilities

    • Generate full-length songs from text prompts, with verses, choruses, bridges, intros, and outros that form a coherent structure.
    • Create expressive, emotionally nuanced vocals in multiple languages, with strict adherence to specified lyrics and bracketed cues (for example [guitar solo]).
    • Produce timed lyrics synchronized to the audio, returning both an MP3 track and the text lyrics used in the song.
    • Support instrumental-only generation for beats, background tracks, and non-vocal music, controlled via prompt instructions.
    • Accept reference images alongside text to influence mood, aesthetic, or thematic choices in the generated music.
    • Offer separate “clip” and “pro” style experiences, with 30-second snippets for quick ideas and multi-minute tracks for more complete songs.
    • Generate high-fidelity 44.1 kHz stereo audio in standard MP3 format, with WAV available on some Gemini and Flow Music surfaces.
    • Embed Google’s SynthID-style watermarking in generated audio, supporting provenance and responsible AI music use.
  • Use cases

    Use Cases for Lyria 3.5

    For creators and musicians, Lyria 3.5 is ideal for rapid song demos from written concepts. A songwriter can paste draft lyrics, specify genre, BPM, and vocal style, and get a full song with structured verses and choruses for evaluation or iteration, all via the Lyria 3.5 API or Gemini surfaces. Example: “Alt-pop ballad, 78 BPM, piano and strings, melancholic mood, [Verse] lyrics pasted below, female vocal, 2-minute song for demo.”

    Marketers and content teams can use Lyria 3.5 to generate branded background music with or without vocals, aligning mood and tempo to campaign assets. For instance: “Upbeat electronic track, 120 BPM, bright synths, no vocals, 60-second loop for product launch video, inspirational mood.”

    Developers integrating music into apps via each::labs and the Gemini-based Lyria 3.5 API can dynamically create soundtracks based on user actions or story states. Example: “Epic orchestral track, 110 BPM, key of D minor, tension-building, instrumental only, 90 seconds for game boss battle.”

    Designers and multimedia artists can pair image mood boards with text prompts so Lyria 3.5 generates audio that matches visual tone, using image inputs alongside detailed music instructions. Example: “Ambient electronic score inspired by attached cityscape images, 95 BPM, soft pads and subtle percussion, 2-minute track, instrumental only.”

  • Tips & tricks

    Tips and Tricks

    Effective prompting is essential for Lyria 3.5. Community and early-usage guidance suggests writing prompts that explicitly specify genre, exact instrument names, numeric BPM, musical key, mood descriptors, and section-level tags or timestamps, in that order. Including section markers like “[Verse]”, “[Chorus]”, “[Bridge]”, or explicit time spans helps Lyria 3.5 maintain structural coherence and align vocals, lyrics, and instrumentation over the full track. You can also guide length by mentioning “30-second loop”, “2-minute song”, or “up to 3 minutes” in the prompt. For tightly controlled vocal behavior, specify language, gender, and tone (for example “soft female vocal in English with intimate delivery”), and clearly state whether you want “instrumental only” or “full vocals with lyrics”.

    Example prompts include: “Dreamy indie pop, 100 BPM, key of G major, warm electric guitars and synth pads, nostalgic mood, [Verse] female vocal in English, [Chorus] bigger harmonies, full-length song around 2 minutes, lyrics about city lights at night.” “Lo-fi hip hop, 85 BPM, dusty vinyl crackle, Fender Rhodes, boom-bap drums, nostalgic mood, instrumental only, 60-second loop for background study music.” “Bossa nova fused with modern R&B, 92 BPM, nylon guitar and electric piano, intimate mood, [Verse] smooth male vocal in English, [Chorus] breathy female vocal in French, bridge instrumental solo, ~3-minute song.”

  • Technical spec

    Technical Specifications

    • Provider / Family: Google DeepMind, Gemini / Lyria music-generation family.
    • Input types: Text prompts and up to ~10 reference images, depending on the surface.
    • Output types: Stereo audio (default MP3, optional WAV in some Gemini / Flow Music contexts) plus generated lyrics text.
    • Sample rate & fidelity: 44.1 kHz high-fidelity stereo audio for full tracks.
    • Max duration: Typically “a couple of minutes” up to around 3 minutes per track, duration steered by the prompt; a clip variant can produce fixed ~30-second snippets.
    • Structure: Verses, choruses, bridges, intros, outros, and timed lyric sections.
    • Token / context: Text+image input with large token limits via Gemini API (listed as 131,072 input tokens).
    • Processing time: Typically under a couple of minutes for full songs in consumer and API workflows (exact latency varies by surface and load).
  • Things to be aware of

    Things to Be Aware Of

    Lyria 3.5 is currently positioned as a preview-style generation workflow, with single-turn song creation rather than fine-grained section editing. If the chorus or bridge does not fit, you generally need to regenerate the entire track; you cannot yet surgically replace one section in-place. Seed control is not exposed, so reproducing an identical result later is difficult, which matters for workflows that rely on deterministic outputs. Official documentation and early reviews also note that quality can vary by genre and prompt detail, and Google encourages users to double-check that the generated track matches their creative intent before publishing. Finally, SynthID watermarking is permanent, so every Lyria 3.5 track remains identifiable as AI-generated, which is important for compliance and rights management.

  • Key considerations

    Key Considerations

    Lyria 3.5 is optimized for end-to-end song generation from prompts, not fine-grained post-editing, so most workflows should treat it as a “generate-and-review” tool rather than a traditional DAW. Users should come with clear genre, mood, BPM, and structure ideas, because detailed prompts significantly improve vocal realism, lyric alignment, and arrangement control. The model works best when you want full songs from text, images, or lyric drafts, and when you are comfortable regenerating entire tracks if a section needs revision. For cost and performance, Lyria 3.5 is exposed in the Gemini API as a preview-style generation path, often priced per full song and optimized for short-to-mid-length pieces rather than long-form albums.

  • Limitations

    Limitations

    Lyria 3.5 does not function as a full digital audio workstation: it cannot perform multi-step mixing, mastering, or detailed track-by-track editing inside the model itself. Fine-grained control over stems, section replacement, or long-form compositions beyond roughly three minutes is limited compared with specialized production tools. Seed-based reproducibility is unavailable, and batch or flex inference options are listed as unsupported in the Gemini API documentation, which restricts some high-scale or deterministic workloads. Pricing for direct Lyria 3.5 API usage is still evolving, with some sources noting per-song costs and others highlighting a lack of fully transparent, long-term pricing. For now, Lyria 3.5 is best viewed as a powerful, vocal-first generator that still requires external tools for deep production work.

Related models

4 models
* FAQ

About Lyria 3.5

01 / 03

What can I create with Lyria 3.5?

You can generate full-length songs from text prompts, including structured verses, choruses, bridges, vocals, lyrics, and instrumental arrangements.