Interested in VLM training with synthetic data? Wondering how to unlock fine-grained visual representations from captions?
I'm happy to present our work FLAIR @CVPR from 10:30 to 12:30 on 15th June, ExHall D Poster #368.
#CVPR2025
Huge credit to my wonderful friends & collaborators @SanghwanKim27, @xyongqin, @zeynepakata, Stephan Alaniz. And thanks to everyone I talked with at both conferences, it was great meeting you. See you at the next one!
Happy to have presented our work "FINER: MLLMs Hallucinate under Fine-grained Negative Queries" at #CVPR2026 as an oral (award candidate) and at #ICVSS2026
[Paper]: https://t.co/6z2H6W2mDw
[Code]: https://t.co/l5t6Vxq30H
We're fortunate (or maybe not) to live in an age where model capabilities advance this fast. Yet I still believe fine-grained alignment between vision and language is still not fully there — which would explain a lot of the problems we still see, and leaves room for future rs.
Excited to be in San Diego at #NeurIPS2025 this week. We'll be presenting Noise Hypernetworks at Hall C,D,E #3605 Thu 4 Dec 11am!
If you're interested in Diffusion/Flow Matching or Reward Alignment, I'd love to chat! Feel free to DM me if you'd like to connect.
Honored to receive one of the Google PhD Fellowships 2025 in ML Foundations for my research on "The Role of the Source Distribution in Generative Modeling: Beyond Fixed Gaussian Noise" :)
Grateful to @zeynepakata, Alexey, @TU_Muenchen, @Googleorg, and all my collaborators!
🎓PhD Spotlight: Karsten Roth
Celebrate @confusezius, who defended his PhD on June 24th summa cum laude!
Karsten has been an @ELLISforEurope and IMPRS-IS PhD student since May 2021, supervised by both @zeynepakata and @OriolVinyalsML. His research has been centered around robust and effective deployment of (large) neural networks in the real world, with particular focus on a method and data-centric perspective for :
🦾 (Multimodal) model pretraining
🧠 Model generalization, reuse and transferability research
🎆 Continual (multimodal) training of such models.
During his PhD, Karsten interned @AIatMeta and @GoogleDeepMind, working on generalization in representation learning and large-scale multimodal pretraining techniques.
🏁 His next stop: @GoogleDeepMind in Zurich!
Join us in celebrating Karsten's achievements and wishing him the best for his future endeavors! 🥳
👇Checkout his selected publications in top-tier conferences such as NeurIPS, ICLR, CVPR or ICCV:
I'm happy to see the potential where we could unlock scalable & cunstomizable visual representations from generative data. Looking forward to further discussion at CVPR.
Interested in VLM training with synthetic data? Wondering how to unlock fine-grained visual representations from captions?
I'm happy to present our work FLAIR @CVPR from 10:30 to 12:30 on 15th June, ExHall D Poster #368.
#CVPR2025
6/
In terms of quantitative experiments, FLAIR show competative results in zero-shot standard, fine-grained, and long image-text retrieval, as well as decent segmentation performance. In zero-shot image classification, FLAIR lags behind CLIP-like models trained on billion-scale.