How to teach AI assistants to infer human minds and proactively provide online assistance?
Sharing our ICML 2026 paper π§ πππ£ππππ§π€: Learning Online Mental Reasoning with Zero Annotations!
π€ Code and weights are fully open source
π https://t.co/xC2yZ1WxWy
My amazing collaborators @ryanlu1219 & @zhuycho11 will be presenting MindZero at the poster session in an hour!
Swing by the poster (Hall A #412) and say hi if youβd like to learn more!
MindZero: Learning Online Mental Reasoning With Zero Annotations
Morning poster session on the first day of the main conference: Tue, Jul 7, 2026, 10:30 AM β 12:15 PM KST, at Hall A #412
https://t.co/wyAgOUP34f
https://t.co/lhVb1fNnGm
One of the challenges for training Theory of Mind models is the lack of ground-truth mental state annotations. How can we address this? It turns out that you don't need any annotations! Introducing MindZero, a self-supervised RL approach for training LLMs to internalize Bayesian inverse planning with ZERO annotations, accepted to #ICML2026.
Key idea: Bayesian inverse planning provides self-supervised reward for training an LLM; after training, LLM generates amortized inferences end-to-end. This is a significant step toward scaling Bayesian inverse planning style Theory of Mind reasoning. Check out @ShunchiZhang's thread for details π
How to teach AI assistants to infer human minds and proactively provide online assistance?
Sharing our ICML 2026 paper π§ πππ£ππππ§π€: Learning Online Mental Reasoning with Zero Annotations!
π€ Code and weights are fully open source
π https://t.co/xC2yZ1WxWy
In summary, we prove that mental reasoning can be effectively learned as a self-supervised skill.
Check out paper and code for details!
π https://t.co/DZMxutVXMl
π» https://t.co/5cNWIwTvwb
Iβll be at #NeurIPS from Dec 2 - 7 and present our spotlight paper πAutoToMπ this Thursday at Poster #2203.
Happy to chat on agent modeling and human-AI collaboration!
π€― Think better visuals mean better world models? Think again.
π₯ Surprise: Agents donβt need eye candyβ they need wins.
Meet World-in-World, the first open benchmark that ranks world models by closed-loop task success, not pixels.
We uncover 3 shocks:
1οΈβ£ Visuals β utility
2οΈβ£ Action data > bigger models
3οΈβ£ Scaling test-time compute = more success
π€ https://t.co/OXn4WfnuTU
π https://t.co/AKRgXhSCJV
π https://t.co/izyjaKTHgO
https://t.co/hd6F9VPGQ2
We spent an additional three months refining and making exciting updates to AutoToM. Here's a summaryβ¨:
1. In addition to achieving SOTA performance on five benchmarks, we conducted further experiments showing that
(a) AutoToM produces human-like confidence estimates as observed in cognitive studies, and
(b) AutoToM enables online mental inference to support embodied decision-making.
This aligns with our long-term vision of developing human-like reasoning and ToM-aware planning.
2. We changed the paper title. The new title reflects our focus on scaling mental inference across different contexts, modalities, mental variables, numbers of agents, and domains. This is achieved through automated agent modeling, which discovers agent models to capture different agents' behaviors and mental states.
3. We evaluated large reasoning models (e.g., o3-mini, Gemini 2.0 Flash Thinking) and found that AutoToM can outperform them in both benchmark performance and embodied decision-making.
π: https://t.co/0hZ4p1SSoZ
State-of-the-art LLMs and VLMs underperform humans by more than 40% on a simplified four-class classification task.
Letβs discuss at https://t.co/RyUUD26m2K!
Many thanks to @Bishop_Gorov, @LemaoLiu, @jieeijjie, @ttchungc and all collaborators!!
π§΅(5/5)
Thanks for sharing our work!
Introducing PhysiCo, a strong successor to the ARC-AGI benchmark for Physical Concept Understanding!
Project page: https://t.co/ekSjcMBA4U (Fully open-sourced, to appear at NAACL 2025)
π§΅(1/n)
Our benchmark covers over 50 physical concepts. Each concept includes multiple phenomena (e.g., gravity includes falling objects, parabolas, and planetary orbits), resulting in 600 total cases.
View our dataset at
https://t.co/VUEc0qEpCg
π§΅(4/n)