The original AC was contributed by @priy2201 in 2018 and it was implemented in a very clever (in retrospect, too clever) way. This AC is what we call "reentrant AC". It is a special autograd function that calls forward AND backward (reentrant!) inside its backward impl.
Reminder about the ✨MARCH 25✨ submission deadline for the Responsible GenAI Workshop @CVPR!!
We welcome 4-page papers at the intersection of responsible AI and generative AI. See full list of topics here: https://t.co/BCOTYxOZE8
I’m very excited to share our work on Gemini today! Gemini is a family of multimodal models that demonstrate really strong capabilities across the image, audio, video, and text domains. Our most-capable model, Gemini Ultra, advances the state of the art in 30 of 32 benchmarks, including 10 of 12 popular text and reasoning benchmarks, 9 of 9 image understanding benchmarks, 6 of 6 video understanding benchmarks, and 5 of 5 speech recognition and speech translation benchmarks. Gemini Ultra is the first model to achieve human-expert performance on MMLU across 57 subjects with a score above 90%. It also achieves a new state-of-the-art score of 62.4% on the new MMMU multimodal reasoning benchmark, outperforming the previous best model by more than 5 percentage points.
Gemini was built by an awesome team of people from @GoogleDeepMind, @GoogleResearch, and elsewhere at @Google, and is one of the largest science and engineering efforts we’ve ever undertaken. As one of the two overall technical leads of the Gemini effort, along with my colleague @OriolVinyalsML, I am incredibly proud of the whole team, and we’re so excited to be sharing our work with you today!
There’s quite a lot of different material about Gemini available, starting with:
Main blog post: https://t.co/NzSycJl7aE
60-page technical report authored by th Gemini Team: https://t.co/CEdMRyYSLo
In this thread, I’ll walk you through some of the highlights.
Introducing the Perception Test, a new multimodal benchmark using real-world videos to help evaluate the perception capabilities of a model: https://t.co/x3zvbwXST8 1/
I learned a lot from my first #FAccT22 session this morning! Thanks to the authors — @priy2201, @ninamarkl, @noagarciad, Yusuke Hirota, and their teams — for the thought provoking papers & discussion!
"Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision"
from FAIR Paris & NYC
Large-scale experiment with 10 billion param RegNet pre-trained by SSL with SwAV on 1 billion random public Instagram photos.
https://t.co/ABSXKFSgqs
Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision
abs: https://t.co/UbTr2ZH0XA
10B parameters dense model, outperforms sota models (supervised and self-supervised) trained on ImageNet on 20 out 25 image classification tasks
Self-supervised learning is really pushing the boundaries of what's possible with deep learning these days. This new paper showcases some of those applications; from improving visual representations to better model robustness and generalization. https://t.co/dpBPlQknS4
We are making model accessible to public after ensuring safety and privacy - we conducted adversarial attacks on the model to the training images can not be recovered. Checkout: https://t.co/hDOr9T8B5A for the model weights, model documentation and model license. (6/6)
Very excited to share our recent work on SEER (SElf-supERvised) which is an order of magnitude bigger and now has 10 Billion dense parameters and is trained on 1 billion randomly selected images from Instagram. (1/6)
We’re pleased to announce new advances in SEER, Meta AI’s groundbreaking self-supervised #computervision model. SEER is now not only much more powerful, it also produces fairer, more robust computer vision models. Learn more:
https://t.co/OferzHr9Ic
The model performance is validated on 50+ computer vision benchmarks including fine-grained classification, geo-localization, copy detection etc. Achieves SOTA self-supervised performance on majority of benchmarks. As model size increases, performance continuously increases.(5/6)
Insightful post from @priy2201 showing that self-supervised learning can allow us to avoid the western-bias of labelled datasets https://t.co/rkDo7qt7Gn
1/ This post from @priy2201 shows SEER, our computer vision system trained using random unlabeled images, outperforming conventional systems in recognizing household items from India, China and Nepal. The reason why is very cool…
https://t.co/ubgV5Abb8N
Many #computervision systems can recognize the item on the left but not the one on the right. Using self-supervised learning, our new SEER model overcomes this limitation, which can help us create systems that work well for everyone around the globe. https://t.co/CzRBaUAg1x
💫 There has been an explosion of interest in self-supervised learning (SSL) for vision and NLP.
In this week's newsletter, we show you recent uses of SSL across various ML tasks from medical image analysis to musical style transfer! Read on below:
https://t.co/iFqITpgnHm