Excited about this new work we've just released, StyleDrop, stylizing text-to-image generation from very few examples (in many cases just one!). Check out the project page for some beautiful results: https://t.co/bwGQbcz3nf
StyleDrop: Text-to-Image Generation in Any Style
introduce StyleDrop, a method that enables the synthesis of images that faithfully follow a specific style using a text-to-image model. The proposed method is extremely versatile and captures nuances and details of a user-provided style, such as color schemes, shading, design patterns, and local and global effects. It efficiently learns a new style by fine-tuning very few trainable parameters (less than 1% of total model parameters) and improving the quality via iterative training with either human or automated feedback. Better yet, StyleDrop is able to deliver impressive results even when the user supplies only a single image that specifies the desired style. An extensive study shows that, for the task of style tuning text-to-image models, StyleDrop implemented on Muse convincingly outperforms other methods, including DreamBooth and textual inversion on Imagen or Stable Diffusion.
paper page: https://t.co/oE5CLhi9Nx
Our short film Dear Upstairs Neighbors is previewing at @sundancefest. 🎬
It’s a story about noisy neighbors, but behind the scenes, it’s about solving a huge challenge in generative AI: control.
Developed by Pixar alumni, an Academy Award winner, researchers, and engineers, here’s how it came together. 🎨
We’re at @sundancefest previewing our new animated short, "Dear Upstairs Neighbors" 📽️
In creating this film, our @GoogleDeepMind team of Pixar alumni, an Academy Award winner, researchers, and engineers designed new AI capabilities specifically for filmmakers.
These tools gave director Connie He a new level of artistic control, allowing her to tell a story she's always wanted to share.
We just dropped Nano Banana Pro, built on Gemini 3. 🍌
With state-of-the-art text rendering, vast world knowledge and studio-quality creative controls, Gemini 3 Pro Image can create and edit more complex visuals, infographics and more. Here’s what’s under the hood. 🧵
This is Gemini 3: our most intelligent model that helps you learn, build and plan anything.
It comes with state-of-the-art reasoning capabilities, world-leading multimodal understanding, and enables new agentic coding experiences. 🧵
Veo is getting new precision editing capabilities that let you easily add or remove elements from a scene - all while preserving the integrity of your original video. 🎥
🍌🍌It's finally here! In addition to the largest ELO lead in lmarena history, I'm most excited about the fact that people really loved using the model. QPS was way above what we expected, and the model racked up 2.5M votes (also a record)! Amazing job team banana 🚀🚀🍌🍌
WoZ at the Sphere was a tour de force of AI for creative industries, combining state-of-the-art super-resolution, outpainting generative video and computer vision research at @GoogleDeepMind to bring it all to life.
What if you could not only watch a generated video, but explore it too? 🌐
Genie 3 is our groundbreaking world model that creates interactive, playable environments from a single text prompt.
From photorealistic landscapes to fantasy realms, the possibilities are endless. 🧵
Veo3 is out and it has a voice! 🗣️
https://t.co/oZWguiIIu9
Veo3 can generate video and audio, including sound effects, ambient noise, even dialogue! The future of AI video is sounding amazing 🤩
#Veo3#AI#GoogleAI#VideoGeneration
Excited to introduce our new Veo 2 capabilities!
Now with reference powered video generation (including style!), camera controls, outpainting, object add/removal & many more:
https://t.co/9t74QkaRc3
Also presenting Flow, our new AI filmmaking tool.
https://t.co/ZHVdxr0zqS
Imagen 4 delivers visuals that pop with richer details, more nuanced color, and better text outputs.
Everyone can make images for free in the Gemini App today: https://t.co/awhPeHZIqm
#GoogleIO
Video, meet audio. 🎥🤝🔊
With Veo 3, our new state-of-the-art generative video model, you can add soundtracks to clips you make.
Create talking characters, include sound effects, and more while developing videos in a range of cinematic styles. 🧵
Introducing Generative Omnimatte:
A method for decomposing a video into complete layers, including objects and their associated effects (e.g., shadows, reflections).
It enables many cool applications, such as video stylization, compositions, moment retiming, and object removal.
What happens when you train a video generation model to be conditioned on motion?
Turns out you can perform "motion prompting," just like you might prompt an LLM! Doing so enables many different capabilities. Here’s a few examples – check out this thread 🧵 for more results!
Motion is the new (and better) language for conditioned video generation, and Daniel @dangengdg shows that its vocabulary should be formed as points and tracks! You can now motion-prompt the same model for object and camera control, motion transfer, and many more!
Excited to introduce our new paper, Generative Omnimatte: Learning to Decompose Video into Layers, with the amazing team at Google DeepMind!
Our method decomposes a video into complete layers, including objects and their associated effects (e.g., shadows, reflections).
Working on layered video decomposition for a few years now, I'm super excited to share these results!
Casual videos to *fully visible* RGBA layers, even under significant occlusions!
Kudos @YaoChihLee, @erika_lu_, Sarah Rumbley, @GeyerMichal, @jbhuang0604, and @forrestercole
🎉Excited to introduce our new paper: Unbounded: A Generative Game of Character Life Simulation!
We build a game of character life simulation that is fully encapsulated in generative models.
🌟We achieve this with:
▶️ A specialized, distilled LLM that dynamically generates game mechanics, narratives, and character interactions in real-time.
▶️ A dynamic regional IP-Adapter for vision models that ensures consistent yet flexible visual generation of a character across multiple environments.
🧵