Analysis
Analysis

Google Makes Gemini Omni 1.1 Flash and Gemini 3.5 Transcribe Production-Ready

Published Aug 29, 2026 Sources checked Aug 29, 2026

Google has moved Gemini Omni 1.1 Flash video generation and Gemini 3.5 Transcribe speech recognition to general availability, expanding production-ready media APIs for developers.

Two specialized media APIs move to general availability

Google's late-August 2026 Gemini API updates make two different media workloads production-ready rather than leaving them in preview.

On August 27, Google released gemini-omni-1.1-flash to general availability. Despite the Omni name, this release is focused on fast conversational video generation and editing. Google says the production model can extend an existing clip, interpolate between first and last frames, and control output resolution from 360p through 4K, with 1080p and 4K produced through upscaling. Google also says the earlier gemini-omni-flash-preview endpoint is scheduled for deprecation on September 30, 2026.

A day earlier, Google made two dedicated speech-to-text models generally available: gemini-3.5-transcribe for non-streaming transcription and gemini-3.5-transcribe-live for low-latency streaming through the Live API. Google's release notes list language detection across more than 85 languages, speaker diarization, word-level timestamps and custom vocabulary biasing for the non-streaming model, while the live model supports interim and finalized events over WebSockets.

Why this matters for AI product architecture

These launches illustrate a broader shift from treating a single general-purpose model as the answer to every multimodal task. Video creation and speech transcription have different latency, state, quality and delivery requirements, so Google is exposing specialized production endpoints that can sit beside general reasoning models.

For developers, the immediate consequence is architectural rather than just benchmark-driven: teams can choose a dedicated transcription path for voice applications and a separate video-generation/editing path for creative workflows, while keeping general Gemini models for reasoning and orchestration. The GA labels also reduce one form of deployment risk compared with preview-only endpoints, although production teams still need to test quality, latency, quotas, regional availability and cost for their own workloads.

Google's current documentation is the source for the capabilities and lifecycle dates above. No independent performance superiority is implied here, and this article does not infer pricing or service guarantees that Google did not state in the checked release notes.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books