Excited about new work on ICL comparing current implicit approaches to explicit task latent variable inference using transformers. Joint work with @EricElmoznino, Leo Gagnon, @sangnie, @dhanya_sridhar and @g_lajoie_ Preprint out now at: https://t.co/0MIgK7r06t
What is the best way of switching between controllers in Reset Free RL?
Read this thread (and come by our poster this Friday morning at ICLR, Hall B) to find out!
#ICLR2024
1/n
🚨 Announcing new paper!!!
TaskMet: Task-Driven Metric Learning for Model Learning
Published at @NeurIPSConf#NeurIPS2023, come check it out if around.
tldr: How to train prediction models tailored to downstream tasks
Work done at FAIR @AIatMeta during my AI residency, with amazing collaborators @brandondamos , @RickyTQChen and @mukadammh
arxiv - https://t.co/7NaQgpgHQY
code - https://t.co/kBKOP6zESe
1/9
Leveraging Unpaired Data for Vision-Language Generative Models via Cycle Consistency
paper page: https://t.co/JAqcmRORps
Current vision-language generative models rely on expansive corpora of paired image-text data to attain optimal performance and generalization capabilities. However, automatically collecting such data (e.g. via large-scale web scraping) leads to low quality and poor image-text correlation, while human annotation is more accurate but requires significant manual effort and expense. We introduce ITIT (InTegrating Image Text): an innovative training paradigm grounded in the concept of cycle consistency which allows vision-language training on unpaired image and text data. ITIT is comprised of a joint image-text encoder with disjoint image and text decoders that enable bidirectional image-to-text and text-to-image generation in a single framework. During training, ITIT leverages a small set of paired image-text data to ensure its output matches the input reasonably well in both directions. Simultaneously, the model is also trained on much larger datasets containing only images or texts. This is achieved by enforcing cycle consistency between the original unpaired samples and the cycle-generated counterparts. For instance, it generates a caption for a given input image and then uses the caption to create an output image, and enforces similarity between the input and output images. Our experiments show that ITIT with unpaired datasets exhibits similar scaling behavior as using high-quality paired data. We demonstrate image generation and captioning performance on par with state-of-the-art text-to-image and image-to-text models with orders of magnitude fewer (only 3M) paired image-text data.
I am grateful to have a helping hand in @pierrelux, @riashatislam and @pierthodo when I first started my research career as an MSc student in 2018. The advice they gave me on reading research papers, organizing my thoughts, and thinking like a researcher is still with me. (1/N)
👩💻💪Encouraging women to have access to mentorship is essential for fostering diversity and inclusivity in ML. Join the @WiMLworkshop breakout session on "Role of Mentorship and Networking" at #ICML2023.
🗓️ 28th July, 11-12
📍 Room 326 A
We will have sure-to-be illuminating roundtable discussions on the importance and how-to's of mentorship. Come find us at Room 326A on 28 July, 11am-12pm. #ICML2023
Late to the party but excited to share that my first paper with @GoogleAI Residency was accepted to #NeurIPS2020: https://t.co/tBZHQX2m7G! We show that an unsupervised Perceptual Quality Metric built using principles from neuroscience works better than supervised models.
Our metric PIM is built using probabilistic representations that encode temporally persistent visual information in videos, learnt with an information theoretic objective. Remarkably, it predicts human ratings across different datasets well, despite never seeing any human labels.
Popular metrics like MS-SSIM are bad at dealing with small pixel shifts and deep metrics like LPIPS have classifier invariances not ideal for quality assessment. (Rows show images with the same quality, according to a metric, as ImageNet-C distortions in the top row)