Gemini Omni 1.1 Flash is our newest multimodal model for video generation and editing. It delivers a new suite of creative capabilities and controls for developers ๐ฅ
With this update you can:
๐ฌ Extend your scenes
๐ฏ Specify starting and ending frames of a shot
โ Add video input references
โจ Upscale your favorite takes up to 4K
โก Test ideas quickly in 360p
See these in action ๐งต
Almost exactly two years ago, I sat down and wrote a 1-pager on what seemed like a wild idea: "Effortless Capture." Under the title, I wrote โJust wave the phone around, and the camera will capture great assets for you.โ The core premise seemed so natural: could we remove the friction from smartphone photography?
Weโve all experienced the frustration of fumbling with settings, missing the shot because we didnโt switch modes fast enough, or feeling like we needed faster fingers to capture a fleeting moment.
The goal was to build a system thatโs intelligent and trustworthy enough that you could just point and let it do the heavy lifting. The camera should just get out of the way of your experiencing the moment.
It was incredibly gratifying to see that vision come to life on the @madebygoogle stage yesterday as Magic Capture, thanks to the efforts of hundreds of Googlers.
I'm thrilled it's finally out in the world, and I can't wait for you all to experience what we've built.
With deep gratitude and admiration to: Krish Desai Hossein Talebi Mauricio Delbracio Francois Bleibel Adi Zicher Anna Lieb Nikhil Karnad Stanley C. Basak Oztas Sivan Doveh Lillian Chen Josรฉ Ricardo Lima Tuo Wang Isaac Reynolds Archan Mehta Ameya Deshpande Chuanlong Xia Marius Renn YiChang Shih Wei (Alex) Hong Sebastian Rodriguez Katherine Raquel Jerome Poichet Dylan Zika Neena Maldikar Shenaz Zack Mistry Robin Dua Brandon Ruffin Michael Milne Soumyadip Ghosh Srimon Chatterjee and many many more!
Very proud of our team. This feature deploys a model that is both the largest image-2-image model we've ever put in Pixel; and also the first diffusion model weโve ever run inside the Pixel Camera.
We're hiring student researchers to work on automated perceptual quality assessment.
Role: Develop & evaluate metrics of human-perceived quality of SOTA generative models, especially for image enhancement.
Requirements: Strong skills in imaging, computer vision, multi-modal LLMs, Python and relevant libraries. Enrolled in a full-time grad program in CS, AI, or related field.
Qualified candidates can send their resume via DM to @hossTale and @2ptmvd.
Our team @ Google is hiring ๐ฅ๐ฒ๐๐ฒ๐ฎ๐ฟ๐ฐ๐ต ๐ฆ๐ฐ๐ถ๐ฒ๐ป๐๐ถ๐๐๐ to develop cutting-edge image & video enhancement solutions. Focus is on research & application of generative models to real imaging problems
Qualified candidates only (incl. new PhDs):
pls DM me/@2ptmvd with CV
To Summarize:
* Pixel: If taken at high zoom (15-30x), you can apply Zoom Enhance with or without cropping to get more details and/or go well past 30x
* Pixel or not: If shot is decent quality & high res already (~12MP), deep cropping with ZE will improve the quality vs just cropping
* Any shot, old or new, good or bad quality - as long as it is around 1MP - can be enhanced fully or cropped and Zoom Enhanced.
* So experiment with different pics and different crops of pics. Even in the same picture, you can get a diversity of results if you change the crop a little.
That's the fun of experimenting with generative models - now in the palm of your hand!
8/n
Here's a little guide on how to make the best use of Zoom Enhance (ZE), our recently released feature on Google Pixel phones. Two broad scenarios:
First: shots of well-lit scenes with good contrast help ZE produce its best work. (It's not yet meant for low-light use)
1/n
Did you ever take a photo & wish you'd zoomed in more or framed better? When this happens, we just crop.
Now there's a better way: Zoom Enhance -a new feature my team just shipped on Pixel. Available in Google Photos under Tools, it enhances both zoomed & un-zoomed images
1/n
My team is hiring a Research Scientist (recent PhD grad or with a couple of years' experience)
Mission: develop fundamental, state of the art tech at the intersection of imaging, vision, and machine learning. Come work w/ @2ptmvd, @hossTale & me
Apply: https://t.co/xUYIqkxddQ
Congratulations to Hossein Talebi & Peyman Milanfar for winning the IEEE Signal Processing Society's Best Paper Award for their paper titled "NIMA: Neural Image Assessment"! #ICASSP2024 https://t.co/t256kLTYLx
Thanks to @IEEEsps for this great recognition, and congratulations to my co-author @hossTale who deserves all the credit.
Paper: https://t.co/buqysmVTpN
We often assume bigger generative models are better. But when practical image generation is limited by compute budget is this still true? Answer is no
By looking at latent diffusion models across different scales our paper sheds light on the quality vs model size tradeoffs
1/5