Whatโs left w/ foundation models? We found that they still can't ground modular concepts across domains.
We present Logic-Enhanced FMs:๐คFMs & neuro-symbolic concept learners. We learn abstractions of concepts like โleftโ across domains & do domain-independent reasoning w/ LLMs.
During long-horizon task planning, a robot must decide what to do, while ensuring each action is geometrically feasible.
To this end, we propose ๐ฝ๐น๐ฎ๐ป๐ป๐ถ๐ป๐ด ๐๐ถ๐๐ต ๐ถ๐ป๐๐ฒ๐ฟ๐น๐ฒ๐ฎ๐๐ฒ๐ฑ ๐๐ถ๐๐ถ๐ผ๐ป-๐น๐ฎ๐ป๐ด๐๐ฎ๐ด๐ฒ ๐๐ต๐ผ๐๐ด๐ต๐๐.
The Visual Concepts Workshop is coming soon to @CVPR 2026!
Join us on June 4 for an exciting discussion on all aspects of visual concepts: from discovery and reasoning, to controllable generation, robotics, and more!
Hope to see you in Denver! #CVPR
Human perception is inherently situated โ we understand the world relative to our own body, viewpoint, and motion.
To deploy multimodal foundation models in embodied settings, we ask:
โCan these models reason in the same observer-centric way?โ
We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness:
- 786 real world egocentric videos
- 2,071 human-annotated QA pairs
Across all tasks, we evaluate 24 state-of-the-art MFMs:
๏ฟฝ๏ฟฝ Best model: 53.9%
๐ง Humans: 91.6%
Models systematically:
โ Confuse head rotation with physical movement
โ Collapse under multi-turn trajectories
โ Fail to maintain persistent world-state memory
๐ We see that maintaining a stable observer-centric representation remains challenging.
As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction.
We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
Happy to bring back the 3rd Workshop on Visual Concepts at @CVPR 2026! Call for papers is now open.
We welcome submissions on the following topics. See our website for more info: https://t.co/LKJxxO5ZGQ
Join us in Denver!
We'll be presenting Deep Schema Grounding at @iclr_conf ๐ธ๐ฌ on Thursday (session 1 #98). Come chat about abstract visual concepts, structured decomposition, & what makes a maze a maze!
& test your models on our challenging Visual Abstractions Benchmark: https://t.co/UnMIoJHN8a
What makes a maze look like a maze?
Humans can reason about infinitely many instantiations of mazesโmade of candy canes, sticks, icing, yarn, etc.
But VLMs often struggle to make sense of such visual abstractions.
We improve VLMs' ability to interpret these abstract concepts.
State classification of objects and their relations (e.g. the cup is next to the plate) is core to many tasks like robot planning and manipulation.
But dynamic real-world environments often require models to generalize to novel predicates from few examples.
We present PHIER, a method that leverages predicate hierarchies to effectively generalize in few-shot scenarios.
Excited to bring back the 2nd Workshop on Visual Concepts at @CVPR 2025, this time with a call for papers!
We welcome submissions on the following topics. See our website for more info:
https://t.co/gk0NgYAcEx
Join us & a fantastic lineup of speakers in Tennessee!
We're organizing the first Workshop on Visual Concepts at @eccvconf with an incredible lineup of speakers!
Join us on Sep 30 afternoon at Suite 3, MiCo Milano ๐ฎ๐น
https://t.co/lnyaci4orM
Work done with the wonderful @maojiayuan, Josh Tenenbaum, @noahdgoodman, and @jiajunwu_cs!
Paper: https://t.co/bvLFJEldnF
Website: https://t.co/zP3dpzkufL
What makes a maze look like a maze?
Humans can reason about infinitely many instantiations of mazesโmade of candy canes, sticks, icing, yarn, etc.
But VLMs often struggle to make sense of such visual abstractions.
We improve VLMs' ability to interpret these abstract concepts.
DSG improves GPT-4V by 5.4 percent point overall on VAD (โ 8.3% relative improvement), and by 6.7 percent point (โ 11.0% relative improvement) on questions that involve counting.
We believe DSG is a promising step towards human-aligned understanding of visual abstractions!
Will be presenting LARC at @CVPR next Thursday morning (session #3) with our fantastic intern @chunfeng3364, as well as @Weiyu_Liu_ and @jiajunwu_cs!
Paper: https://t.co/n8WgZ0mmwu
Website: https://t.co/L3HhfGwZrh
Code: https://t.co/dHIN7fCpxv
How can we learn 3D visual grounding with natural supervisionโonly looking at QA pairs, without ground truth bounding boxes or object classification labels?
We inject explicit language priors, e.g., the symmetric property that A near B โ B near A, in structured vision models.
LARC shows strong data efficiency and generalization to new concepts & dataset, and is a promising step towards injecting structured visual reasoning frameworks with explicit language-based priors, for learning in settings without dense visual supervision.