New preprint from @ghazalkhn and @raghavlite!
The dominant paradigm for multimodal document retrieval has been to represent pdf pages as image inputs to a VLM. We show that this approach is suboptimal for text-rich scientific documents.
With the same VLM, keeping text as tokens and interleaving with figure images leads to better performance than page-level images -- despite not being trained on such sequences!
Text embeddings alone also handle figure-based queries quite well -- by using captions and in-text references.
Overall, text-based embeddings are still quite effective for multimodal retrieval! Lots more interesting findings in the paper: https://t.co/OL2U9adNTn
It was really fun to meet new people and discuss agent environments. Thanks to the workshop organizers for putting together such a great event! Here is the slide deck from the talk: https://t.co/uyNh7l0vIn
👋 Hello #SanDiego!
We are excited to present our latest work bridging Causal Inference and GenAI at #NeurIPS2025.
For tackling unobserved confounders, come visit on Dec 5, 2025 at 11 AM PST.
📍Poster 2516
Coupling Generative Modeling and an Autoencoder with the Causal Bridge🧵👇
📢 New preprint announcement!!
- Most image encoders are trained independently before being integrated into a VLM, resulting to generic, query-agnostic image representations that constrain downstream VLM performance.
- 🚀 We introduce the Text-Guided Semantic Image Encoder (TIE), which produces query/image-conditioned image representations, enabling VLMs to operate on task-relevant image features.
- Across nine image-to-text benchmarks, TIE-based VLMs achieve +1.5–1.3 average improvements, with gains of up to +6 points on DocVQA and InfoVQA.
- Notably, despite using only half the image tokens, and therefore offering faster inference, TIE-based VLMs outperform their standard counterparts.
Paper: https://t.co/MS1ZPg1A4a
Excited to be in Montreal this week for COLM 2025!
I have two papers accepted this year:
1️⃣ The Unlearning Mirage, a dynamic framework for evaluating LLM unlearning.
2️⃣ With @Harsh_N_Lalai10, using 20 Questions as a creative framework to systematically evaluate geographical bias in LLMs.
Check out the individual threads!
📢 New Preprint from @raghavlite on Multimodal Contrastive Learning: Breaking the Batch Barrier (B3) 📢
TL;DR: Smart batch mining based on community detection achieves state of the art on the MMEB benchmark.
Preprint: https://t.co/BU85rYwDBp
Code: https://t.co/bo2xLRBG7Q
📢 New Preprint from @raghavlite on Multimodal Contrastive Learning: Breaking the Batch Barrier (B3) 📢
TL;DR: Smart batch mining based on community detection achieves state of the art on the MMEB benchmark.
Preprint: https://t.co/BU85rYwDBp
Code: https://t.co/bo2xLRBG7Q
I will be presenting ✨GenEOL: Harnessing the Generative Power of LLMs for Training-Free
Sentence Embeddings✨at #NAACL2025!
We propose a simple and effective training-free inference time method to achive high quality embeddings
- Beats other trainig-free methods by upto 4 points on STS benchmark.
- Notable gains over training-free, unsupervised methods on the MTEB benchmark
- Robust to prompt perturbations
Link: https://t.co/Txq10epNBR
📅 4/30, Wed, 2:00-3:30 Hall 3
Happy to chat more about embeddings and my current research!
I just published Tips for Writing NLP Papers https://t.co/u9zbHdsMC5
I wrote it for my students so I don't have to sound like a broken record (and edit papers for the same issues over and over again 😃). But some of you might find it useful too.
For very long, OntoNotes has served as the most important benchmark for coreference resolution.
We are happy to present "LongtoNotes" that extends OntoNotes to up to 8X longer documents.
Paper : https://t.co/HZlvTP8ZzK