📢New Preprint 📢
Do VLMs reason fairly over text and image inputs? Absolutely not! When given conflicting image-text pairs, they strongly favor either text or image based on the task, and this bias is highly linked to sample difficulty. See our paper:
https://t.co/Sz8IgNeeqz 1/n