🚀 We’ll be presenting PerceptionRubrics at tomorrow’s poster session!
🕓 Time: July 9, Thursday, 10:30 AM–12:15 PM KST
📍 Location: Hall A, #312
📖 Topics: MLLM Dense Perception, Human-aligned Evaluation, Rubrics
🔗 Paper: https://t.co/Lcm0pqPTjz
Come by and chat with us!
🚀 Is MLLM perception evaluation really saturated? Check out our ICML 2026 paper!
While scores are getting harder to distinguish, models still make unacceptable visual mistakes in real-world use. Even one critical visual mistake can make the whole response unreliable to humans.
👉 We introduce PerceptionRubrics, a rubric-based evaluation framework for multimodal perception, which shows strong agreement with human Elo scores from VisionArena.
Our benchmark includes 1,038 information-dense images and 10K+ atomic rubrics across 7 domains: natural scenes, OCR, GUIs, charts, STEM, logic puzzles, and creative/cultural images.
We evaluate 20+ mainstream MLLMs, including GPT-5.5. Key findings:
1️⃣ Models often pass fragmented atomic checks but fail strict conjunctive perception, revealing a clear Reliability Gap.
2️⃣ Perceptual reliability remains especially challenging in information-dense domains such as GUIs, documents, and structured data.
3️⃣ As an automated benchmark, PerceptionRubrics shows strong alignment with human preference.
🏠 Project Page: https://t.co/kNjXaxf8Ym
📊 Code & Data: https://t.co/CjjrIUda7C
📄 Paper: https://t.co/sWf37vnz2C
Welcome model evaluations on PerceptionRubrics and would love to hear your feedback!
#ICML2026 #MLLM #MultimodalAI #ComputerVision
🚀 Is MLLM perception evaluation really saturated? Check out our ICML 2026 paper!
While scores are getting harder to distinguish, models still make unacceptable visual mistakes in real-world use. Even one critical visual mistake can make the whole response unreliable to humans.
👉 We introduce PerceptionRubrics, a rubric-based evaluation framework for multimodal perception, which shows strong agreement with human Elo scores from VisionArena.
Our benchmark includes 1,038 information-dense images and 10K+ atomic rubrics across 7 domains: natural scenes, OCR, GUIs, charts, STEM, logic puzzles, and creative/cultural images.
We evaluate 20+ mainstream MLLMs, including GPT-5.5. Key findings:
1️⃣ Models often pass fragmented atomic checks but fail strict conjunctive perception, revealing a clear Reliability Gap.
2️⃣ Perceptual reliability remains especially challenging in information-dense domains such as GUIs, documents, and structured data.
3️⃣ As an automated benchmark, PerceptionRubrics shows strong alignment with human preference.
🏠 Project Page: https://t.co/kNjXaxf8Ym
📊 Code & Data: https://t.co/CjjrIUda7C
📄 Paper: https://t.co/sWf37vnz2C
Welcome model evaluations on PerceptionRubrics and would love to hear your feedback!
#ICML2026 #MLLM #MultimodalAI #ComputerVision
Want to ask your humanoid to get novel objects you want from a novel table? We introduce HERO, first achieving open-vocabulary visual loco-manipulation using human language queries! Now we can ask the humanoid to see and grasp the target object (eg, a carrot instead of a tissue box) with whole-body coordination.
How do we do it? Key designs and findings in this work:
We propose a residual-aware end-effector tracking policy that tracks the target accurately in a closed loop.
We find that the robot's forward kinematics are quite inaccurate, and thus we propose residual neural forward models that correct the end-effector FK and base leg odometry.
We design a modular system powered by visual foundation models that achieves generalizable grasping capabilities with 83% success rate on novel daily objects and daily scenes.
Check out our new project:
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation
Project page: https://t.co/4QtbYne1I2
Paper: https://t.co/FLEjaGkRpW
Huge congratulations to @jieneng_chen and collaborators on being selected for an oral presentation (top 6% of submissions!) -- very well deserved!
This is exciting work, and I hope it gets the attention it truly deserves.
🚀 Open Vision Reasoner and Perception-R1 are both accepted at #NeurIPS2025!
Also releasing OVR’s Cold Start data with rich reasoning here:
🔗 https://t.co/nahzrxCj6K
I’m at 🌴NeurIPS this week! I’ll present our 🤔 Open Vision Reasoner tomorrow afternoon (4:30 pm & #1907). Feel free to stop by and say hi!
🕓Time: 4:30pm-7:30pm, December 5
📍Location: Exhibit Hall C, D, E, #1907
📖Main content: MLLM & LLM reasoning, Cold Start & RL, Cognitive Behaviors.
🚀 Open Vision Reasoner and Perception-R1 are both accepted at #NeurIPS2025!
Also releasing OVR’s Cold Start data with rich reasoning here:
🔗 https://t.co/nahzrxCj6K
🔥 Thrilled to release our new multimodal RL work: Open Vision Reasoner!
A powerful 7B model with SOTA performance on language & vision reasoning benchmarks, trained with nearly 1K steps of multimodal RL.
Our journey begins with a central question:
Can the cognitive behaviors of LLMs transfer to MLLMs for advanced visual reasoning?
Yes—here's how👇
🤖 What’s inside:
👉 Two-stage training: Cold-start (text-only) + massive multimodal RL (on Qwen2.5-VL-7B)
👉 SOTA results: MATH500 95.3%, MathVision 51.8%, MathVerse 54.6%
👉 In-depth analysis: How cognitive behaviors emerge, transfer & scale up
💡 3 key insights on behavior transfer:
1️⃣ Behavior transfer emerges surprisingly early in cold start due to linguistic mental imagery.
2️⃣ Cold start broadly memorizes visual behaviors, while RL critically discerns and scales up effective patterns.
3️⃣ Transfer strategically favors high-utility behaviors such as visual reflection.
📈 More highlights:
– Reward & response length co-scale with sequence expansion
– Cold start impairs perception, while RL enhances.
– Perceptual RL faces temporary unscalability, where rewards grow but response length remains stagnant.
🌐 Project: https://t.co/p0Uqmo3KAc
🐙 Code: https://t.co/7t7hl5Bcve
📄 Paper: https://t.co/SAkcQeY5Jb
Huge thanks to the amazing team—this strong 7B reasoner, the insightful analyses, thorough evaluations, and elegant presentation wouldn���t have been possible without you all.
Thanks for reading! Feel free to chat if you have any questions.
Why can language teach MLLMs to see?
Read thoughts on:
👀 Visual priors embeded in language
➡️ How reasoning transfer across modalities
🧠 How SVGs from #Gemini 3 pro hint at a model’s inner “cognitive imaginery”.
✨ Blog: https://t.co/XgA4ZwKV0A (Stay tuned!)
🤯 Think better visuals mean better world models? Think again.
💥 Surprise: Agents don’t need eye candy— they need wins.
Meet World-in-World, the first open benchmark that ranks world models by closed-loop task success, not pixels.
We uncover 3 shocks:
1️⃣ Visuals ≠ utility
2️⃣ Action data > bigger models
3️⃣ Scaling test-time compute = more success
🤗 https://t.co/OXn4WfnuTU
🌍 https://t.co/AKRgXhSCJV
📄 https://t.co/izyjaKTHgO
https://t.co/hd6F9VPGQ2
Perception in Reflection is here at ICML! 🎉
Come explore how VLMs can enhance perception through reflection—stop by and chat!
📅 When: July 17, 4:30–7:00 PM (that’s today!)
📍 Where: East Exhibition Hall A-B, #E-2509
🚀 Check out our ✈️ ICML 2025 work: Perception in Reflection!
A reasonable perception paradigm for LVLMs should be iterative rather than a single-pass.
💡 Key Ideas
👉 Builds a perception-feedback loop through a curated visual reflection dataset.
👉 Utilizes Reflective Unlikelihood Training to capture preferences and prevent behavioral collapse. 👉 Proposes a policy-critic inference architecture (Reflective Perception), enabling perception & reflection separately in multi-turn dialogues.
🙌 Highlights
🎯 RePer gradually shifts visual attention toward human-aligned regions
🧩 Reflective Perceptual Learning acts as free-form preference optimization, unifying DPO, LiPO, and enabling fine-grained supervision via explicit feedback
🧵 RePer is an accurate, informative captioner with significantly fewer hallucinations
✍️Arxiv: https://t.co/eBYTNNiist
🐼Github: https://t.co/SwcmkYufne
💻Project Page: https://t.co/ZTFeEaypH0
🪞 We'll present Perception in Reflection at ICML this week!
We introduce RePer, a dual-model framework that improves visual understanding through reflection.
Better captions, fewer hallucinations, stronger alignment.
📄 https://t.co/DMH3PeLY1A
#ICML2025@yanawei_@JHUCompSci
🚀 Open Vision Reasoner (OVR)
Transferring linguistic cognitive behaviors to visual reasoning via large-scale multimodal RL.
SOTA on MATH500 (95.3%), MathVision, and MathVerse.
💻 Code: https://t.co/eB82Pawlue
🌐 Project: https://t.co/tMAFw1oe3N
#LLM@yanawei@HopkinsEngineer
Just three words and one link—beautiful project pages in surprising styles! ✨
Try magic Anycoder by @_akhaliq here 👉 https://t.co/UCHKvXI1TF
Original pages have totally different vibes! (https://t.co/ZTFeEaypH0)
🥳 Thanks @_akhaliq for featuring our OVR!
We’re continuously iterating on both models and data to release even more 🚀 powerful versions—stay tuned!
Check out the nearly 1K-step multimodal RL and in-depth cognitive behavior analysis here:
✍️ ArXiv: https://t.co/5Y0DUDGOTl
🐼 GitHub: https://t.co/JDuXXlIfOP
🤩 Project page: https://t.co/nEUk2IMegz
I’ll also be at ICML after the 16th and present our work Perception in Reflection (https://t.co/ZTFeEayXwy) on 17th—feel free to stop by and chat!