I wish research papers have a section in the appendix that is titled - "What did not work".
Although the main paper should outline "what works", it's worth writing about the series of failed experiments.
Introducing a New Era of Personalized AI Art! 🚀
I'm excited to share our latest AI model that can transform your selfies into stunning works of art, preserving your unique facial features. 🎨✨
Introducing a novel zero-shot image-to-image model designed for personalized and stylized portraits. Learn how it both accurately preserves the similarity of the input facial image and faithfully applies the artistic style specified in the text prompt →https://t.co/ULxXfeOJl3
Google announces Imagen 3
discuss: https://t.co/SrjnFQhR5K
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of evaluation. In addition, we discuss issues around safety and representation, as well as methods we used to minimize the potential harm of our models.
Here's my take on the Sora technical report, with a good dose of speculation that could be totally off. First of all, really appreciate the team for sharing helpful insights and design decisions – Sora is incredible and is set to transform the video generation community.
What we have learned so far:
- Architecture: Sora is built on our diffusion transformer (DiT) model (published in ICCV 2023) — it's a diffusion model with a transformer backbone, in short:
DiT = [VAE encoder + ViT + DDPM + VAE decoder].
According to the report, it seems there are not much additional bells and whistles.
- "Video compressor network": Looks like it's just a VAE but trained on raw video data. Tokenization probably plays a significant role in getting good temporal consistency. By the way, VAE is a ConvNet, so DiT technically is a hybrid model ;) (1/n)
Our recent work on "Accelerating Batch Active Learning Using Continual Learning Techniques" will appear in #ICML2023 workshop on "Data-centric Machine Learning Research". Joint work with @arnaved and Jeff Bilmes. Summary 🧵 soon!
I would like to express my immense gratitude to my parents, Ujwala and Nandkishor, for their unconditional love, faith, support and motivation. I cannot thank you both enough for all the sacrifices you have made for me, but I dedicate my dissertation to you.
Congratulations Dr. @surajkothawade for successfully defending his thesis! Suraj is the first student (advised by me at @UT_Dallas) who has defended his Ph.D.!!
@surajkothawade will be joining @Google next month!
I could not have undertaken this journey without my dissertation committee members: Prof. @lakshman_tamil, Prof. Gopal Gupta, Prof. @ganramkr and Prof. Yu Xiang. I am sincerely thankful to them for generously providing their knowledge and expertise.
📢Trying to improve your object detection models but see them consistently failing on the same kind of scenarios? #ECCV2022
We present TALISMAN: Targeted Active Learning for Object Detection with Rare Classes and Slices using Submodular Mutual Information, accepted to ECCV 2022!
Happy to announce that the code for our #ECCV2022 paper "TALISMAN: Targeted Active Learning for Object Detection with Rare Classes and Slices using Submodular Mutual Information" is now available at: https://t.co/u3FDuN59Zu
Check it out 🚀
📢 Trying to tackle class imbalance and out-of-distribution(OOD) scenarios in medical imaging? #MICCAI2022
Obtaining labeled medical data is difficult and expensive since it requires expert annotators like doctors; such imperfections in data exponentiate this difficulty.
1/n
📰CLINICAL and DIAGNOSE papers are available as a part of MICCAI proceedings here: https://t.co/XVZaB2Cf5P
This is joint work with my fantastic collaborators: Akshit Shrivastava, Atharv Savarkar, Venkat Iyer, Prof. Ganesh Ramakrishnan and my amazing advisor Prof. @rishiyer.
4/n