{"id":44756,"date":"2026-05-19T23:45:41","date_gmt":"2026-05-19T23:45:41","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/44756\/"},"modified":"2026-05-19T23:45:41","modified_gmt":"2026-05-19T23:45:41","slug":"googles-gemini-omni-turns-images-audio-and-text-into-video-and-thats-just-the-start","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/44756\/","title":{"rendered":"Google&#8217;s Gemini Omni turns images, audio, and text into video \u2014 and that&#8217;s just the start"},"content":{"rendered":"<p id=\"speakable-summary\" class=\"wp-block-paragraph\">When Google launched <a href=\"https:\/\/blog.google\/innovation-and-ai\/technology\/ai\/google-gemini-ai\/#performance\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Gemini three years ago<\/a>, the goal was to build a multimodal large language model \u2014 a single neural network that was trained on text, image, audio, and video and could generate content in any of those formats.<\/p>\n<p class=\"wp-block-paragraph\">Today, <a href=\"https:\/\/techcrunch.com\/events\/google-io\/\" rel=\"nofollow noopener\" target=\"_blank\">at its Google I\/O developer conference<\/a>, the company took a concrete step toward that goal with Gemini Omni, a new family of multimodal models that Google CEO Sundar Pichai says will be able to \u201ccreate anything from any input.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Omni will start with video. Users can now combine images, audio, video, and text, and rather than simply stitching those inputs together, Omni reasons across all of them to produce a consistent output. The result is high-quality videos that reflect an understanding of physics, culture, history, and science.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Omni also lets users edit photos with plain text commands rather than complex editing software, similar to <a href=\"https:\/\/techcrunch.com\/2026\/02\/26\/google-launches-nano-banana-2-model-with-faster-image-generation\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Google\u2019s Nano Banana<\/a>.<\/p>\n<p class=\"wp-block-paragraph\">Google already has a dedicated video model, <a href=\"https:\/\/techcrunch.com\/2025\/10\/15\/google-releases-veo-3-1-adds-it-to-flow-video-editor\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Veo<\/a>, that lets users turn text and images into videos, and even <a href=\"https:\/\/techcrunch.com\/2026\/04\/02\/google-now-lets-you-direct-avatars-through-prompts-in-its-vids-app\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">direct and customize avatars<\/a>. But Google DeepMind director of product management Nicole Brichtova says that today\u2019s release is more than a Veo update: \u201cIt\u2019s the next step towards the progression of combining the intelligence of Gemini with the rendering capabilities of our media models.\u201d<\/p>\n<p class=\"wp-block-paragraph\">One example that Koray Kavukcuoglu, DeepMind\u2019s chief technologist, gave reporters during a media briefing on Monday: When Omni was given a simple prompt like \u201ca claymation explainer of protein folding,\u201d it quickly rendered a video of a stop-motion explainer with a voice-over that said, \u201cProteins start as chains of amino acids. They fold into patterns like the alpha helix and flat sections called beta sheets, forming a perfect three-dimensional shape.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The long-term vision for Omni is broader, involving the model being used to do things like generate images from audio, or audio from video.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cWhen we first announced Gemini, it was our first AI model to be natively multimodal,\u201d Pichai said during the briefing. \u201cWe knew that training it on a combination of text, code, audio, images, and video would give it a deeper understanding of the world. With world models, AI is moving from predicting text to simulating reality. Gemini Omni is the next step in that direction.\u201d<\/p>\n<p class=\"wp-block-paragraph\">As part of the release, users will also be able to create videos with their own digital avatars \u2014 something OpenAI popularized on its now-defunct Sora app with Cameos. To prevent deepfakes, users will have to go through a dedicated product onboarding, which involves recording themselves and speaking out a series of numbers, per Brichtova. The avatar then gets stored for future use.<\/p>\n<p class=\"wp-block-paragraph\">Additionally, all videos created with Omni will include Google\u2019s SynthID digital watermark, which allows users to verify if videos were generated via the Gemini products.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The first model in the family is Gemini Omni Flash, which will roll out today to the Gemini app, YouTube Shorts, and AI creative studio Flow. Flash will be capable of rendering 10 seconds of video, which Brichtova says isn\u2019t a model limitation, but rather a decision based both on a desire to get it into more hands and an anticipation that most users won\u2019t want to make much longer videos yet. Longer video durations are in the pipeline for the near future, though.<\/p>\n<p class=\"wp-block-paragraph\">Google seems to be pitching Omni Flash as more of a consumer tool. The examples Brichtova and Gabe Barth-Maron, a research engineer at DeepMind, gave on a call with TechCrunch of uses for digital avatars were all personal: Making a video of yourself winning an award or going to the moon, or removing a passerby from the background of a video you took on vacation.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Barth-Maron put it more simply: \u201cThey\u2019re like personalized memes.\u201d<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe definitely did focus on making this easy to use for consumers,\u201d Brichtova said. \u201cNot many video models have breached that chasm with consumers, so this is our play to do that.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The ease of use comes with a caveat: Brichtova and Barth-Maron noted that editing prompts will need to be highly specific, otherwise Omni risks over-editing or unintentionally altering elements the user wanted to keep \u2014 a problem Nano Banana users would have run into.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" height=\"383\" width=\"680\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/Omni-Flash-ball.gif\" alt=\"\" class=\"wp-image-3123989\"\/>Image Credits:Google<\/p>\n<p class=\"wp-block-paragraph\">Despite the near-term consumer focus, Omni\u2019s enterprise and <a href=\"https:\/\/techcrunch.com\/2026\/01\/13\/googles-update-for-veo-3-1-lets-users-create-vertical-videos-through-reference-images\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">creative implications<\/a> are obvious, and Google will make Omni available via API in the coming weeks. The avatar-generating tool \u2014 a capability that is\u00a0available today on Shorts \u2014 is something Google expects content creators to pick up. But more broadly, an end-to-end multimodal workflow could be transformative for advertisers and filmmakers.<\/p>\n<p class=\"wp-block-paragraph\">Startup Luma AI is building something similar, <a href=\"https:\/\/techcrunch.com\/2026\/03\/05\/exclusive-luma-launches-creative-ai-agents-powered-by-its-new-unified-intelligence-models\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">an agentic tool<\/a> that can generate an entire ad campaign based on a short brief and a product image, powered by its own \u201cunified\u201d model.<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe\u2019re actually pretty proud of the model\u2019s text-rendering capabilities, which is really useful for things like advertising,\u201d Brichtova said. \u201cIf you want a product somewhere, or even just a slogan, it needs to be accurate\u00a0\u2026 We definitely anticipate filmmakers and other kinds of creators are going to be using this model as well.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The more professional use cases might be better served by the Omni Pro model, which should perform better across all Omni tasks. Google hasn\u2019t said when it will release Pro yet, but Brichtova said that will happen when \u201cwe feel like we\u2019re at a point where we have a step change above Flash.\u201d<\/p>\n<p>Catch up on the rest of Google IO 2026\u2019s big news<\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/techcrunch.com\/2026\/05\/19\/google-search-as-you-know-it-is-over\/\" rel=\"nofollow noopener\" target=\"_blank\">Google Search as you know it is over<\/a><\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/techcrunch.com\/2026\/05\/19\/google-updates-its-gemini-app-to-take-on-chatgpt-and-claude\/\" rel=\"nofollow noopener\" target=\"_blank\">Google updates Gemini app to take on ChatGPT and Claude<\/a><\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/techcrunch.com\/2026\/05\/19\/google-introduces-gemini-spark-a-24-7-agentic-assistant-with-gmail-integration\/\" rel=\"nofollow noopener\" target=\"_blank\">Google introduces Gemini Spark, a 24\/7 agent assistant with Gmail integration<\/a><\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/techcrunch.com\/2026\/05\/19\/how-to-use-googles-new-information-agents\/\" rel=\"nofollow noopener\" target=\"_blank\">How to use Google\u2019s new information agents<\/a><\/p>\n<p>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" rel=\"nofollow noopener\" target=\"_blank\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/p>\n","protected":false},"excerpt":{"rendered":"When Google launched Gemini three years ago, the goal was to build a multimodal large language model \u2014&hellip;\n","protected":false},"author":2,"featured_media":44757,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9],"tags":[5044,27120,132,7543,27013,22922,22923,15189],"class_list":["post-44756","post","type-post","status-publish","format-standard","has-post-thumbnail","category-google","tag-deepmind","tag-gemini-omni-flash","tag-google","tag-google-deepmind","tag-google-gemini-omni","tag-google-io","tag-google-io-2026","tag-veo"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/44756","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=44756"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/44756\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/44757"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=44756"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=44756"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=44756"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}