🔥 Excited to share my first work from CLVR Lab!
We introduce N2M a module that guides the robot to the suitable pose for executing the manipulation policy. Check out this post!
🚀 Introducing N2M!
In mobile manipulation, the performance of a manipulation policy is very sensitive to the robot’s initial pose. N2M guides the robot to a suitable pose for executing the manipulation policy.
N2M comes with 5 key features - Check them out in the posts below!
🔥 Excited to share my first work from CLVR Lab!
We introduce N2M a module that guides the robot to the suitable pose for executing the manipulation policy. Check out this post!
Temporal Cortex (TC) shouted out by Jensen Huang! @nvidia@NVIDIAAI
TC is the best bio-inspired memory for AI agents: one tht doesnt rot.
For hackers who want long-running AI agent memory, TC is a MUST.
Sign up here: https://t.co/Og9eBm8lH0
@2_9sigma@im_sihun@ksunw0209@nurijunsu
1/ We just released π0.7 — a steerable generalist robot model with emergent capabilities.
I want to share a bit of the backstory, because π0.7 taught me something surprising about where robot learning is heading. A thread on bittersweet lessons 🧵
🧠 What if large models could read each other’s minds?
Our new paper (#neurips2025 spotlight), “Thought Communication in Multiagent Collaboration”, explores how large model agents can share latent thoughts, not just messages.
📷https://t.co/IzX1Hy7K0K (CMU × Meta AI × MBZUAI)
Imagine teams of agents that don’t just talk, but directly read each other’s minds during collaboration. That’s the essence of Thought Communication, which goes beyond the fundamental limits of natural language, or any observed modalities.
🧩 Theoretically, we prove that in a general nonparametric setting, both shared and private latent thoughts can be identified from model states under a sparsity regularization.
Our theory ensures these recovered representations reflect the true internal process of agent reasoning, and that the causal structure between agents and their thoughts can be reliably recovered.
⚙️ Practically, we introduce ThoughtComm, a general framework for latent thought communication. Guided by the theory, we implement a sparsity-regularized autoencoder to extract thoughts from model states and infer which are shared or private.
This lets agents not only know what others are thinking, but also which thoughts they mutually hold or keep private — a step toward real collective intelligence.
Across diverse models, communication beyond language directly enhances coordination, reasoning, and collaboration among LLM agents.
🔮 In line with recent studies, we believe this work further highlights the importance of the hidden world underlying foundation models, where understanding thought, not just observational behavior, becomes central to intelligence.
#MultiAgent #LLMs #Causality #AI #ML
Joint work with @zhuokaiz @ Zijian Li @xie_yaqi @ Mingze Gao @LizhuZhang@kunkzhang
What if robots could improve themselves by learning from their own failures in the real-world?
Introducing 𝗣𝗟𝗗 (𝗣𝗿𝗼𝗯𝗲, 𝗟𝗲𝗮𝗿𝗻, 𝗗𝗶𝘀𝘁𝗶𝗹𝗹) — a recipe that enables Vision-Language-Action (VLA) models to self-improve for high-precision manipulation tasks.
PLD couples real-world residual reinforcement learning with standard supervised fine-tuning — letting robots discover, recover, and distill their own data flywheel.
Quick 🧵
Rollouts in the real world are slow and expensive. What if we could rollout trajectories entirely inside a world model (WM)?
Introducing 🚀Ctrl-World🚀, a generative manipulation WM that can interact with advanced VLA policy in imagination. 🧵1/6
How can we help *any* image-input policy generalize better?
👉 Meet PEEK 🤖 — a framework that uses VLMs to decide *where* to look and *what* to do, so downstream policies — from ACT, 3D-DA, or even π₀ — generalize more effectively!
🧵
I presented our work "Grounded Vision-Language Interpreter for Integrated Task and Motion Planning" as a poster at the SAFE-ROL Workshop in #CoRL2025!
Thanks to everyone who stopped by and checked out our bimanual cooking demo video!
If you're around CoRL, would love to chat!
Come see what robot learning can do for surgical automation!
We’re excited to host the first Workshop on Automating Robotic Surgery with an amazing lineup of speakers.
🗓️ Sept. 27 09:30AM - 12:30PM
📍 Floor 3F, Room E7
🌐 https://t.co/Lp10McSvb2
#CoRL2025#CoRL@corl_conf
Official results are in - Gemini achieved gold-medal level in the International Mathematical Olympiad! 🏆 An advanced version was able to solve 5 out of 6 problems. Incredible progress - huge congrats to @lmthang and the team! https://t.co/pp9bXF7rVj
We can synchronize multiple diffusion models for one collaborative generation. But WHY does it work? HOW should we synchronize?
**SyncSDE** explains
- WHY it works: probabilistic framework for synchronizing diffusion models
- WHERE to focus: where synchronization heuristics should be strategically focused, to reduce trial-and-error
- Strong results on diverse tasks
Accepted in CVPR 2025!!
Arxiv: https://t.co/tQi5bLT37G
Website: https://t.co/GgE6sovWzA
[1/n]