skimmed through S1 by Skild AI and it makes me less convinced this is just retrieval/interpolation. if the policy can change the execution plan while preserving intent what exactly is the context inducing inside the policy?
really cool. curious what the context is actually identifying here. is it inferring the task/intent online or mostly retrieving and interpolating within behaviors already learned? and when it goes OOD does the context still help it replan?
The best thing about this robot? It has no hands.
No hands = no manipulation problem.
No grasping, no dexterity nightmares, no trying to teach a robot how to pick up random shit.
Sometimes the best robotics solution is just deleting the hardest problem entirely
Behavioral cloning mystery
https://t.co/VqxzvcfGSx
I wrote a new blog post about "mysteries" in behavioral cloning that appear with real-world robot data (e.g., overfitting is "good"). I also tried to demystify them and shared my thoughts!
I don't date people traning robotics VLA.
Even kissing them feels like end-to-end embodied control. You think he's staring into your eyes. His Vision Encoder is actually running Feature Extraction at 30Hz, estimating your lip Pose in real time. He leans in. Not romance. A K-step Action Chunk.
Hand trajectory: smooth. Tracking Error: minimal. Force Control: absolutely fucked.
Then the lighting changes. He immediately backs off:
"Wait. Distribution shift. This is way OOD from our 500 real-world demos. We need Human-in-the-loop intervention."
But love isn't Imitation Learning. You don't need Teleop Data, Calibration, or an Action Planner to kiss someone.
Sometimes you just lean in. No Prompt. No Action Horizon. No Closed-loop Feedback.
The policy collapses. VRAM explodes. Control loops melt. Because you were never the optimal Trajectory in my training set.
You were the one thing I never wanted to generalize.
Disclaimer: I still love VLA folks haha. Just sharing a poem I saw in Chinese. Epic.
I don't date people traning robotics VLA.
Even kissing them feels like end-to-end embodied control. You think he's staring into your eyes. His Vision Encoder is actually running Feature Extraction at 30Hz, estimating your lip Pose in real time. He leans in. Not romance. A K-step Action Chunk.
Hand trajectory: smooth. Tracking Error: minimal. Force Control: absolutely fucked.
Then the lighting changes. He immediately backs off:
"Wait. Distribution shift. This is way OOD from our 500 real-world demos. We need Human-in-the-loop intervention."
But love isn't Imitation Learning. You don't need Teleop Data, Calibration, or an Action Planner to kiss someone.
Sometimes you just lean in. No Prompt. No Action Horizon. No Closed-loop Feedback.
The policy collapses. VRAM explodes. Control loops melt. Because you were never the optimal Trajectory in my training set.
You were the one thing I never wanted to generalize.
Disclaimer: I still love VLA folks haha. Just sharing a poem I saw in Chinese. Epic.
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours,
• this human data scaling law implied a scaling law on never seen robot data,
• both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge
🧵
🤖 A robot that transforms between humanoid and dexterous hand?
Introducing Handroid, a reconfigurable robot with a shared 27-DoF body 🖐️→🚶→🖐️:
• The same joints, different roles
• Expanded robot task space
• Shared sensing and control interfaces
https://t.co/wCVGuo9MmJ
New work with @nvidia: evaluating robot policies entirely inside a world model. The policy acts, the model imagines the consequences, and the imagined evals predict real-world results. 🧵
real vs world-model rollout side by side📷
As far as I can tell this is the first time pi0.5 has run on Spot. The openpi repo has no quadruped support. There's literally one GitHub discussion asking if it's possible with no answer. Now there is one.
@sama when I use it with memory on, it tends to over prioritize whats saved in memory even when its not relevant to the current chat lk would love more control over when memory is used