Humans learn by interacting with the world: How can we teach robots to similarly adapt to ambiguous dynamics? Introducing Wiggle and Go! #CoRL2026. We find that a specialized model can learn rope properties from just one interaction, and complete complex dynamic tasks.
The robot wiggles a rope once, figures out how it behaves, then goes, using the estimated dynamics to execute the task. No real-world training data. No retries. #robotics #robotlearning
Introducing PointZero—a 3D world model pre-trained without robots.
Dexterous manipulation requires understanding diverse 3D dynamics, but current approaches rely on expensive robot data. So how can we pre-train a 3D dynamics model without it?
PointZero introduces a simple idea: learning to complete 3D point tracks yields transferable 3D dynamics!
Our model achieves SOTA results on:
🔥 Zero-shot 3D dynamics
🔥 Action-conditioned 3D dynamics (post-training)
🔥 Imitation learning (post-training)
Details and links 👇 1/6
Humans learn by interacting with the world: How can we teach robots to similarly adapt to ambiguous dynamics? Introducing Wiggle and Go! #CoRL2026. We find that a specialized model can learn rope properties from just one interaction, and complete complex dynamic tasks.
The robot wiggles a rope once, figures out how it behaves, then goes, using the estimated dynamics to execute the task. No real-world training data. No retries. #robotics #robotlearning
Results in real:
3.55 cm average error on 3D target striking vs 15.29 cm for uninformed baselines
0.95 Pearson correlation between sim and real rope motion on unseen trajectories
World-action models typically imagine the future in RGB – but are pixels really the right representation for robotics?
Our bet is no: RGB spends capacity on fine-grained details and variation that are often irrelevant for robot policies.
DINO features, point tracks, and depth capture more useful features like semantics, motion, and geometry.
But no single modality captures everything — how can we effectively combine them?
We introduce ✨ModAR✨, which predicts the future one modality at a time, with each prediction informing the next. We train from scratch and find that this formulation performs best.
Our 30.1M scratch-trained model even outperforms a 6B video-model-initialized model finetuned on the same data!
🧵 [1/8]
Today we are introducing Atlas! An autoregressive and multimodal DiT for generating image and depth frames. Atlas is a state-of-the-art 3D reconstruction model and outperforms all open-source models. @HaoZhang623 and I had lots of fun pushing the 3D capability of this model; can't wait to see you build with it!
How can we get robot hands to “hear” slip and contact through microphones, and react to them?
We’re excited to share VibeAct, an approach that uses piezoelectric microphones embedded in robot fingertips to estimate contact and slip, then learns reactive policies from this tactile feedback!
https://t.co/7jMLZNpYzp
Let your robots hear slips with A-SLIP! 🤖🎧
How can a robot detect in-hand slip and estimate its direction and magnitude without cameras or fragile tactile skins?
A-SLIP uses piezoelectric microphones embedded in grippers to hear it.
https://t.co/VZWCZ7XaPe
🧵1/7
Excited to share SoftAct, a framework for retargeting human manipulation demos to soft robot hands using explicit contact force reasoning! How do you transfer human skill to a hand that looks and moves nothing like yours🐙🖐️? It turns out VR environments can let us capture privileged force interaction demonstrations to help. 🧵1/7
Learning from human videos often requires restrictive, carefully choreographed human motions.
We propose ✨3PoinTr✨: a scalable way to pretrain from casual human videos. It bridges the embodiment gap by learning 3D scene evolution, enabling learning from natural human motions.
🤖🦾✍️Why is robot grasping hard? We usually blame contacts, kinematics, geometries, perception, and so on. But what if the object is just being spiteful?
🔥We propose a game-theoretic grasp synthesis method as a two-player game between the robot and an adversarial (spiteful) object.
💡In this formulation we achieve SOTA grasp success rates, without training data - just using optimization tools (Augmented Lagrangian + Iterative Best Response).
📑Arxiv: https://t.co/mH2ocISuHy
🌐Website: https://t.co/HKayRv9yRD
🧵[1/7]
[1/7] Teaching dexterous robot hands to perform functional grasps usually needs hours of teleoperation, manual labeling, or pre-scanning object meshes.
Not anymore.
🔥We are excited to introduce Web2Grasp that learns functional multi-finger grasps straight from web images of human hand-object interactions (HOI).
No human demos. No object scans. Just web images.
👉https://t.co/lOcWgMOpu5
@CMU_Robotics@CarnegieMellon