Do VLMs truly reason or memorize data? We show
❗even transformations like rotation & scaling still throw them off
❗finetuning fails to generalize across domains
https://t.co/EIfNtDRby5
📅 Stop by our poster on Wed @ 4:30 PM and
chat with @xavierohan#COLM2026@COLM_conf
🚀 Excited to share our ACL 2026 paper:
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
LVLMs can be confidently wrong when relying on language priors instead of the image. VAUQ checks whether model confidence is truly grounded in visual evidence 👀
Excited to share our new paper: "A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning."
Multimodal LLMs increasingly listen as well as see - so we asked: can you attack them through speech? It turns out you can, easily.
🤖 How do you detect VLA failures during execution with only trajectory-level labels?
Excited to share our new paper:
"Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring"
[1/4] The human eye doesn't process every single pixel of a video continuously—it focuses on what changes.
So why are our video AI models wasting compute on redundant frames?
Introducing Swift Sampling: a test-time technique inspired by the human visual system. 🧠👇
[1/4] 🧠 The Great Modality Test: Are MLLMs Blind, Deaf, or Both?
Humans are perception pros: We instinctively know what information to trust.
Blindfolded, you describe the chirping birds.
Headphones on, you focus entirely on the silent visuals.
We robustly use the available and reliable modality.
Can SOTA MLLMs do the same?
We put the best Multimodal LLMs to the ultimate test.
📢📢 Can we automate how humans instantly perceive if a generated video captures a human action done wrong?
Details: https://t.co/vV2T1EG0zX
Paper: https://t.co/18H8t5niJW
Our key idea is to build a manifold of correctly performed actions from several real-world videos. Next, we measure how far off a generated video is from this manifold to get a robust measurement of action correctness.
We perform extensive subjective, quantitative, and quantitative evaluation of different actions, generative models, and MLLMs.
Work done by @xavierohan Youngsun, Ananya, and Audrey
Happy to share I’ll be presenting our paper at #NeurIPS in San Diego!
🚀 GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
(with Prof. @SharonYixuanLi)
📄 https://t.co/wl1BP4pnlc
📍 Exhibit Hall C/D/E — Booth #1401
🕟 Dec 3 (Wed), 4:30–7:30 PM PST
🎉 Excited to share that our ICML 2025 paper on LLM hallucination detection has been accepted!
Poster📍: East Exhibition Hall A-B #E-2510 — Tue, July 15 | 4:30–7:00 p.m. PDT
Would love to chat and connect — come say hi! 😊
Can coding agents autonomously implement AI research extensions?
We introduce RExBench, a benchmark that tests if a coding agent can implement a novel experiment based on existing research and code.
Finding: Most agents we tested had a low success rate, but there is promise!
Come meet me and my students at #CVPR25 Below is a list of works on interpretability and controllability of generative models that my students will be presenting🪩
1. Revelio: Interpreting and leveraging semantic information in diffusion models https://t.co/J9zL5gVhqn
2. Concept Steerers: Leveraging K-Sparse Autoencoders for Controllable Generations: https://t.co/s38NBWPTTw
3. What's in a latent? Leveraging Diffusion Latent Space for Domain Generalization https://t.co/ybt12yuaa4
4. Improving Physical Object State Representation in Text‑to‑Image Generative Systems: https://t.co/8kNHiWypHQ
5. Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models https://t.co/F5NsP5jufL