
Google has unveiled Gemini Omni Flash, the first model in a new “Omni” family that turns any combination of text, image, audio, and video into a generated clip and lets users refine the result through plain-language editing. The model rolls out today to Google AI Plus, Pro, and Ultra subscribers via the Gemini app and Google Flow, with free access landing in YouTube Shorts and the YouTube Create app starting this week.
Announced at Google I/O by Koray Kavukcuoglu, CTO of Google DeepMind and Chief AI Architect, Omni extends the company’s video generation stack from the text-and-image-driven approach of Veo into a multimodal pipeline where reference media can be mixed in a single prompt. Google frames the model as the point where “Gemini’s ability to reason meets the ability to create,” with image and audio outputs planned to follow video.
For those tracking this category, Omni Flash slots in alongside Veo 3.1, which Google updated earlier this year with native vertical formatting, 4K upscaling, and improved character consistency, and the broader Gemini image lineage that the company has been pushing through its Nano Banana releases. The pitch is different from those earlier tools: rather than treating text-to-video and image-to-video as separate workflows, Omni accepts mixed inputs natively.
Multimodal inputs as the central proposition
Omni takes images, text, audio, and video as references, then blends them into a single generated output. A user can supply a still photograph, a short reference clip for motion or lighting, an audio file for tempo or mood, and a written description, and the model is meant to reconcile all of it into one cohesive video. At launch, the audio input is restricted to voice references, with Google saying other audio input types will follow.
Conversational editing across multiple turns
The second pillar is iterative, language-driven editing. Once a clip exists, users can refine it through successive natural-language instructions: change the environment, the camera angle, the style, or specific details, and the model is supposed to preserve continuity across turns.
Google’s examples include prompts such as, “When the person touches the mirror, make the mirror ripple beautifully like liquid, and the person’s arm turns into reflective mirror material,” with the original scene’s characters and physics carried through subsequent edits. The framing positions Omni as something closer to a conversational compositor than a one-shot generator.
Maintaining a coherent thread across multiple edits has been one of the persistent weak points in this category, and Google is explicitly claiming progress here. The promise of editing video through dialogue with a model, with characters and scene logic intact across iterations, is genuinely compelling if it holds up beyond a handful of turns.
Credit: GooglePhysics, world knowledge, and explainer content
Google is also positioning Omni as a step forward on physical plausibility, pointing to improved handling of gravity, kinetic energy, and fluid dynamics. Demonstration prompts include a violinist playing a piece, a marble rolling through a chain-reaction track, and a claymation-style stop-motion explainer of protein folding.
The model also leans on Gemini’s broader knowledge base, which Google argues lets it generate explainer videos that are factually grounded rather than purely pattern-matched. For documentary and educational filmmakers, that is the more interesting claim, though it is the kind of capability that needs scrutiny on a case-by-case basis. Plausible-looking visuals that are subtly wrong on the facts are arguably worse than obviously synthetic ones.
Digital avatars and the audio question
A more provocative addition is Avatars, which lets users create a digital version of themselves that can generate videos with their own voice. Google says it is still working through how to bring broader audio and speech editing capabilities (specifically, editing audio inside an existing video) to users responsibly, and those features are not yet enabled.
That is a notable contrast with some competitors. The ability to alter dialogue in an existing clip is exactly where AI video tools become genuinely dangerous in editorial and political contexts, and Google’s decision to hold that capability back, at least for now, reflects an awareness of the territory that the industry has just been forced to confront.
SynthID by Google.Watermarking and provenance
Every video generated by Omni carries a SynthID digital watermark, which Google has been embedding across its generative media output, including Veo and Lyria. The watermark is verifiable through the Gemini app, Gemini in Chrome, and Google Search.
Provenance tooling matters more than ever in the wake of the Hollywood-wide backlash earlier this year against ByteDance’s Seedance 2.0, after fabricated clips featuring Tom Cruise, Brad Pitt, and other A-list actors went viral within hours of that model’s release. SAG-AFTRA, the Motion Picture Association, and Disney condemned the model and the practices it enabled. The pressure on every model vendor to ship verifiable watermarking has increased sharply since then, and Google is clearly positioning Omni to be on the right side of that conversation.
Where this lands for working filmmakers
The question for professionals is less about whether Omni can produce a striking clip from a creative prompt, and more about whether it slots into any real production pipeline. Google Flow, the company’s broader creative workspace built around its video models and used in projects like Darren Aronofsky’s Primordial Soup partnership with DeepMind last year, is the integration point for paid subscribers. The strategy of bundling video generation, image generation, and now multimodal editing under one paid umbrella is now fully evident.
For short-form creators on YouTube Shorts and the YouTube Create app, free access changes the math considerably and will almost certainly produce a substantial volume of Omni-generated content in the coming weeks. For narrative and commercial work, the more meaningful test is whether the conversational editing model can maintain consistency long enough to be useful for previs, mood reference, or stylized inserts, areas where AI video tools have started to find a practical footing.
It is also worth noting where Google sits in the competitive landscape. The pace has accelerated dramatically: Kling 3.0 introduced native 4K and multi-shot sequencing in February, Seedance 2.0 hit the market the same month with photorealistic results that triggered the Hollywood reaction described above. Multimodal input handling is one of the few clearly differentiating axes left in this race, and that appears to be where Google is staking its claim with Omni.
Will multimodal inputs actually change how you work with AI video, or is this another impressive demo cycle? Don’t hesitate to let us know in the comments below!