Check out Proactive Co-Creator on
@GoogleAIStudio , a human-AI belief alignment demo I vibe coded: https://t.co/BEcchP0K89
🧠 See & edit the AI's uncertainty via belief graph. It asks clarifying questions before creating!
📷 Try Image ➔ Story ➔ Video. You can even remix it!
Our short film Dear Upstairs Neighbors is previewing at @sundancefest. ���
It’s a story about noisy neighbors, but behind the scenes, it’s about solving a huge challenge in generative AI: control.
Developed by Pixar alumni, an Academy Award winner, researchers, and engineers, here’s how it came together. 🎨
Exciting new work from @sihyun_yu and our team at Google Deep Mind!
Memory-Augmented Latent Transformers (MALT) Diffusion, a new diffusion model specialized for long video generation!
https://t.co/gcDZr5mVbf
Check out our tech report on proactive T2I agents that ask clarification questions to reduce uncertainty!
With this agent, we obtain 2 times higher VQAScore in just 5 turns!🤯 We've open-sourced our agent code powered by #Gemini@GoogleDeepMind! 🚀
Code:https://t.co/67dTGDQpZI
legit treflip by Veo 2 with just one error (wheel inversion in the middle). legs move realistically for a treflip 🤯. skateboarding videos are notoriously hard to generate
Tired of endless prompt tweaking? We've released a tech report on proactive text-to-image agents powered by #Gemini@GoogleDeepMind! Our agents ask clarifying questions and use belief graphs to understand what you really want.
https://t.co/jRwjxtqALx
https://t.co/XFVvva7r96
2/ website: https://t.co/atH5wzRudu
Our approach has two key design decisions. First, we use a causal encoder to compress images and videos in a shared latent space.
We introduce W.A.L.T, a diffusion model for photorealistic video generation. Our model is a transformer trained on image and video generation in a shared latent space. 🧵👇
Have you ever wondered about emergent intelligence in robotic agents?
This work shows interesting emergent intelligence and behaviors in blind navigation agents! Blind agents learn maps as they navigate. This allows them to navigate as successfully as an agent with vision
How do 'map-less' agents navigate? They learn to build implicit maps of their environment in their hidden state!
We study 'blind' AI navigation agents and find the following 🧵
How can we fill in missing pulsative sensor data? Prior state-of-the-art fails in our novel setting, despite its well-defined temporal structure.
Checkout our #NeurIPS2022 paper, PulseImpute, @ 4 pm CST!
arxiv: https://t.co/Hbv2x7ZvkP
github: https://t.co/bTManuLyEH
Dense self-supervised learning from multiple 3D viewpoints → dense feature representations that generalize both to novel object instances and to novel categories of instances.
Checkout our #NeurIPS2022 paper!
arxiv: https://t.co/rxzdrScII1
github: https://t.co/5n98W7Wykt
We model indoor environments using FPV panoramic navigation graphs and introduce a visiolinguistic transformer model, LED-Bert, which scores the alignment between navigation graph nodes and dialogs and achieves SOTA performance on the LED task!
✨Transformer-based Localization from Embodied Dialog with Large-scale Pre-training✨ has been accepted as an oral at @aaclmeeting!
https://t.co/L7SSGfqk0w
w/ @RehgJim
Today, along with my collaborators at @GoogleAI, we announce DreamBooth! It allows a user to generate a subject of choice (pet, object, etc.) in myriad contexts and with text-guided semantic variations! The options are endless. (Thread 👇)
webpage: https://t.co/EDpIyalqiK
1/N
A new year, a new shameless twitter plug: Check out our Toys4K 3D object dataset
4K instances, 105 categories, 15+ instances per category
https://t.co/K85J007iru