Loading article…
Google’s new Gemini Omni AI model generates video from text, audio, and images. Available now for subscribers and free on YouTube Shorts with SynthID.
Google has released Gemini Omni, a multimodal AI model capable of generating and editing video through text, audio, and image inputs, marking a shift toward more interactive, physics-aware content creation [1, 2]. The model, which debuted at Google I/O 2026, is now available to paid Gemini subscribers and will roll out to YouTube Shorts and YouTube Create later this week at no cost to users [2].
| At a glance | |
|---|---|
| Product | Gemini Omni Flash |
| Developer | |
| Primary Input | Text, Audio, Images, Video |
| Availability | Gemini App, Flow, YouTube Shorts |
Gemini Omni represents a departure from previous text-to-video models like Google’s Veo by allowing users to modify existing footage through natural language prompts [1, 2]. The model analyzes multiple input types simultaneously to perform complex edits, such as changing a video's background, lighting, or camera angle without requiring the user to regenerate the entire clip [2, 3]. Google claims the model incorporates an understanding of real-world physics, including gravity and fluid dynamics, to improve the realism of generated sequences [1, 2].
The tool also features an avatar creation system that allows users to generate digital likenesses using their own facial and voice data [1, 3]. To address concerns regarding AI-generated content, Google has embedded all output with SynthID, a digital watermarking technology designed to verify the origin of the media [2, 3]. While the model offers significant creative flexibility, it currently carries a 10-second limit on video generation length, which may constrain its use for longer-form production [3].
The rollout of Gemini Omni Flash serves as a direct expansion of Google’s AI film-making ecosystem, integrating with the company’s existing Flow platform [1, 2]. By offering a free version through YouTube Shorts and YouTube Create, Google is positioning the technology to compete for the attention of casual creators and social media users, rather than just professional editors [2, 3].
Unlike previous iterations that focused on purely AI-generated scenes, Omni is designed to function as a "world model" that can interpret and manipulate real-world context [2]. This capability is intended to bridge the gap between static media and interactive, AI-driven storytelling [2]. However, the company has not yet clarified whether creators will have the ability to restrict others from remixing their content using these new AI tools [1].
The success of Gemini Omni will likely depend on how effectively it balances its advanced editing capabilities with the practical limitations of its current generation length. Whether the model’s physics-based simulations can consistently produce high-quality, realistic output remains the primary test for its adoption among professional and casual creators alike [1, 3].
Coverage is mostly measured — 248 of 260 reports stay neutral.
Every Monday — the token unlocks, Fed dates & catalysts set to move crypto and markets this week. So you’re never blindsided.
Free · 3-min read · one-click unsubscribe
AI-assisted synthesis by the TrendWatcher Editorial Desk · sourced from 3 outlets · Sep 12, 2026 · How we report
As of September 2026, businesses can improve Google AI Overview visibility by implementing structured content like FAQPage schema and maintaining strong Google Business Profile signals. Key signals include having more than 40 reviews, ensuring primary category accuracy, and uploading new photos within the prior 90 days.
As of September 2026, Google AI subscriptions include voice capabilities for Gmail, Docs, and Keep that allow for conversational inbox searches, hands-free document writing, and the conversion of spoken input into structured lists. These features are available to various tiers of Google AI Plus, Pro, and Ultra subscribers.
Google AI Sheets canvas is a tool that turns spreadsheet data into interactive, no-code mini-apps based on natural language prompts. As of August 2026, it allows users to create dynamic, read-write layers over spreadsheet data that update in real time.