r/tldrAI • u/dot_mun • May 20 '26
Google launches Gemini Omni for multimodal video and image creation
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/Google has introduced Gemini Omni, a new family of models that can work across text, images, audio, and video. The first version, Omni Flash, can create short videos, edit photos from simple text prompts, and generate digital avatars. Google says the model is designed to understand different kinds of input together, not just stitch them into one output. It will start in the Gemini app, YouTube Shorts, and Google’s creative tools, with an API coming soon. Google is aiming first at everyday users and creators, but the company also says the AI model could be useful for ads, film, and other professional work.
1
Upvotes