Why do embedding models struggle with binding, a core requirement for multi-object understanding?
We find that objects don't decompose into simple components. Binding is possible, but requires going beyond the linear representation hypothesis.
Co-led with @a_uselis
π§΅π
@a_rubique@a_uselis@coallaoh If you want to know more and discuss, come to our poster at #ICML2026 π°π· next week!
ποΈ Thu, Jul 9 β’ 2:30 PM β 4:15 PM KST
π Hall A #2903
Interestingly, these models that have generalization also have low-complexity binding functions!
So there is a pattern:
CLIP doesn't generalize β has a high-complexity binding function
Models that generalize β have a low-complexity binding function
The object components are responsible for the uni-modal binding observed in prior work.
But if both image embeddings and text embeddings contain the binding information, why is that information not aligned? π€
We use MLPs to approximate input embeddings (CLIP or DINO) from scene descriptions (inspired by @EricElmoznino's work on complexity).
Turns out those approximations encode concepts well.
But they do not generalize to unseen objects. Even after observing 90% of other objects.
Why do embedding models struggle with binding, a core requirement for multi-object understanding?
We find that objects don't decompose into simple components. Binding is possible, but requires going beyond the linear representation hypothesis.
Co-led with @a_uselis
π§΅π
We find that CLIP embeddings have an approximately additive structure: scenes decompose into objects, and objects decompose into concepts.
But the object components are not fully explained by their concepts alone. They contain extra information that may be crucial for binding.
First, we formalize binding.
Scenes contain objects. Objects are combinations of concepts.
A model exhibits binding if it can not only recognize the concepts present in a scene, but also identify which concepts belong together as the same object.