Lyria 3.5 Expands to Gemini and API: Pricing, Evaluation Gaps and Early User Feedback
Google has expanded Lyria 3.5 beyond Flow Music into the Gemini app and Gemini API. Here is what is verified about access, pricing, evaluation methodology, safety and the limits of early feedback.
Google expanded Lyria 3.5 on September 4, 2026 from its earlier Flow Music launch into the Gemini app and the Gemini API. This is a distribution change for an existing model, not a brand-new foundation-model launch: Google originally introduced Lyria 3.5 on July 29, 2026. The distinction matters because many reports are describing the September rollout as if the model itself first appeared this week.
What changed on September 4
Google's September 4 announcement says Lyria 3.5 is now available in the Gemini app and Gemini API, in addition to Google Flow Music, Google AI Studio and Google Vids. In Gemini, users can choose or describe a genre, select vocal or instrumental output, use templates and request shorter or longer tracks. Google says the Gemini rollout is global; the current Gemini music overview adds that music generation is available where the Gemini app is available and requires users to be at least 18.
For developers, the current Gemini API music-generation documentation identifies the full-song model as lyria-3.5. It generates 44.1 kHz stereo audio, accepts text and image inputs, can use up to 10 images with a prompt, returns MP3 by default and can return WAV. The API documentation describes full songs as lasting a couple of minutes with duration influenced by the prompt; Google's broader Lyria product page describes the model family as supporting tracks up to three minutes.
Exact model identity and architecture
Google DeepMind's Lyria 3.5 model card describes Lyria 3.5 as a text-to-music system using latent diffusion over temporal audio latents. The card says it was trained on audio data annotated with text captions at different levels of detail, with deduplication, safety filtering and quality filtering during preprocessing. Training used Google TPUs with JAX and ML Pathways.
This is not a coding or general-purpose reasoning model, so coding benchmarks such as SWE-bench Verified, SWE-bench Pro, Terminal-Bench or tool-use leaderboards are not applicable. Treating a music-generation model as if it should have comparable coding scores would create a meaningless ranking.
Pricing and access
The current Gemini Developer API pricing page lists lyria-3.5 at $0.08 per successful full-song request on the paid API and shows no free API tier for the model. For comparison inside Google's own family, Lyria 3 Clip Preview is listed at $0.04 per 30-second song and the legacy Lyria 3 Pro Preview at $0.08 per full song.
That means the September consumer rollout and developer rollout have different economics. Google says Lyria 3.5 is available to Gemini users globally through the app, while API use is a paid per-request product. The checked official pages do not publish a universal per-user Gemini generation quota, a production latency distribution, or a cost-per-minute figure for Lyria 3.5, so those values should not be inferred.
What Google's evaluation actually proves
The model card says Google used both human and automated evaluations across curated prompts, working with music experts and comparing several music-generation systems including Lyria 2. Evaluation dimensions included music quality and aesthetics, vocal quality, audio fidelity and prompt adherence.
Google reports that Lyria 3.5 improved significantly over Lyria 2 on audio fidelity and that lyric generation showed better prompt adherence. However, the public model card does not provide the evaluation sample size, per-benchmark scores, confidence intervals, evaluator agreement, the identities of every comparison model, or enough raw artifacts to reproduce a numeric leaderboard.
That makes the evidence useful but limited. It supports Google's claim of an internal improvement over Lyria 2, but it does not support a precise public ranking such as “Lyria 3.5 is better than Suno, Udio or every competing music model.” No such cross-vendor conclusion is warranted from the checked sources.
Safety and provenance
Google says all generated Lyria audio carries an imperceptible SynthID watermark. The API documentation also says prompts are filtered, including requests for specific artist voices or copyrighted lyrics. The DeepMind model card describes additional red-teaming, safety reviews, supervised fine-tuning, reinforcement learning from human and critic feedback, dataset filtering and product-level safeguards.
These controls do not answer every rights question. The checked model card says training used audio datasets and explains filtering, but it does not publish a complete itemized training-data inventory. Users planning commercial distribution should therefore rely on the terms that govern the specific Google product or API they use rather than assuming that model availability itself is a blanket commercial-rights grant.
Early public feedback is mixed and anecdotal
Public feedback is still too early and self-selected to treat as a measured consensus. In a September 4 Reddit discussion, one commenter criticized the output quality and price. In a separate September 4 user test, the poster praised acoustic/emotional output but reported weaker results for EDM and faster styles, difficulty outside English, and a recurring synthetic artifact.
Those posts are useful as bug and quality leads, not as benchmark evidence. They involve unknown prompting skill, unknown sampling settings, tiny sample sizes and a highly self-selected audience. I found no directly retrievable X post with enough first-hand detail to score Lyria 3.5 quality reliably, so no X consensus is claimed here.
Practical tradeoffs
Lyria 3.5 is now easier to access than at its July launch, and the API exposes a straightforward per-song price. Its strongest documented product features are full-song structure, vocals and lyrics, prompt-controlled timing, image-conditioned music, multilingual prompting and SynthID provenance.
The biggest evidence gap is not availability but independent measurement. A useful next step would be a public, same-prompt evaluation against other music generators with disclosed prompt sets, blind human preference methodology, language and genre breakdowns, audio-fidelity metrics, latency, failure rates and total cost per usable song. Until that exists, Google's qualitative evaluation and early user reports should remain separate evidence categories rather than being merged into a single ranking.
This article is built from the source material below. Open the originals for full context and the latest updates.