🤖 What if a humanoid robot could make a hamburger from raw ingredients—all the way to your plate?
🔥 Excited to announce ViTacFormer: our new pipeline for next-level dexterous manipulation with active vision + high-resolution touch.
🎯 For the first time ever, we demonstrate ~2.5 minutes of continuous, autonomous control—combining active vision, high-res touch, and high-DoF robot hands SharpaWave — to complete complex, real-world tasks.
Code is fully released; check out our:
Homepage: https://t.co/I8kVJO3XPw
Paper link: https://t.co/BFqpOwXYbT
Github: https://t.co/345eTpTx0y
Humanoid robots shouldn't just follow pre-defined movements—they should perform! 🤖✨
Introducing UniAct: A unified model for multimodal motion generation and action streaming.
Most humanoids are limited to pre-designed moves. UniAct changes the game by allowing robots to generate live action sequences from:
📝 Text instructions 🎶 Music rhythms 📉 Spatial trajectories 🔄 Cross-modal signals
Whether it’s dancing to a beat or following a complex path, UniAct brings humanoid robots to life in real-time. 🚀
🔗 Project: https://t.co/ecAZk8ta8i
📄 Paper: https://t.co/ZfriCthUtz
Excited to see what @SharpaRobotics has achieved recently. Proud that we shaped the ViTacFormer prototype—even before the hand was fully product-ready, we were already making a hamburger (and doing many other tasks).
My takeaway: how you fuse vision + touch matters a lot. With clean, high-quality teleop data, you can get impressive generalization with even ~50 demos.
This is the most dexterous task I’ve seen a humanoid do so far.
Fully autonomous powered by Sharpa’s CraftNet (VTLA) — using tactile feedback to continuously fine-tune the last-millimeter interaction.
@karanjagtiani04 Human data is actually much more scalable than robot data. Large part of our training data are in the wild human data, which i believe is quite scalable.
This might be my "aha moment" of 2025:
With our new robotics foundation model, Large Video Planner, we train a robot planner from large-scale video data. It works so well that we can use it directly for robot planning.
Two moments really blew my mind:
First: right after our model training, I fed in an image of my hand and my MacBook and asked it to close the laptop—when the Apple logo appeared exactly as the lid came down, I couldn’t help but feel impressed (and excited).
Second demo: picking up the brush — check the 3D consistency. Even the brush shadow is remarkably accurate, and it can even infer what the Franka arm (at the corner) should look like.
@bruk_phi We actually use a lot of human dataset, large part of the training dataset are human data, and we also use human hand as an intermediate representation for robot planning.
Introducing Large Video Planner (LVP-14B) — a robot foundation model that actually generalizes. LVP is built on video gen, not VLA. As my final work at @MIT, LVP has all its eval tasks proposed by third parties as a maximum stress test, but it excels!🤗
https://t.co/wjD54YFK3k
Because of the model’s remarkably accurate 3D consistency and human-hand prediction, we can use it directly for robotics planning—and it already pulls off a bunch of amazing dexterous tasks: zero-shot and, in some cases, first-ever.
It also makes me believe even more that human-centric, in-the-wild data has huge potential to further boost robotics pretraining.
Been working on this for a long time—Large Video Planner is finally out. We found that large-scale video pretraining can actually teach models surprisingly accurate physics.
With strong camera consistency and realistic hand shaping, we can use the model directly for robotics planning. When I tried using our dexterous hand to pull off extremely challenging zero-shot tasks—often for the first time—made it clear just how capable this model is.
Introducing Large Video Planner (LVP-14B) — a robot foundation model that actually generalizes. LVP is built on video gen, not VLA. As my final work at @MIT, LVP has all its eval tasks proposed by third parties as a maximum stress test, but it excels!🤗
https://t.co/wjD54YFK3k
🎉🎉🎉 We won the champion in the solo dance contest at the first World Humanoid Robot Games, partner with @UnitreeRobotics ! Here is the full video!
Training the robot to perform a long-term dancing (2:30 mins) with stability, smoothness, and agility is much more challenging than we expected. The robot needs to dance with the rhythm, keep global position, move dynamically and cannot fall.
You cannot cherry pick on the playing field. More technical details will be released in the future.
🤖 What if a humanoid robot could make a hamburger from raw ingredients—all the way to your plate?
🔥 Excited to announce ViTacFormer: our new pipeline for next-level dexterous manipulation with active vision + high-resolution touch.
🎯 For the first time ever, we demonstrate ~2.5 minutes of continuous, autonomous control—combining active vision, high-res touch, and high-DoF robot hands SharpaWave — to complete complex, real-world tasks.
Code is fully released; check out our:
Homepage: https://t.co/I8kVJO3XPw
Paper link: https://t.co/BFqpOwXYbT
Github: https://t.co/345eTpTx0y
Thank you so much, @adcock_brett, for featuring our new work, ViTacFormer!
Generalizable and robust manipulation remains a long-standing and challenging goal in robot learning—we’re excited to keep pushing the boundaries.
More exciting things are on the way—stay tuned! 🚀
@OfficialLoganK UC Berkeley researchers introduced ViTacFormer, a unified visuo-tactile pipeline for robot manipulation
It fuses high-resolution visual and tactile data using cross-attention and enables multi-fingered hands to perform precise, long-horizon tasks
There is a raging debate over sensory modes and redundancy
How much is enough and is sensory overload an issue
Perhaps the key is redundant modes are the fabric that hold actions together to solve long term planning
Think fascia
This work wouldn’t have been possible without the incredible support from great collaborators @cosm_para13983 and @KaifengZhang4, and my amazing advisors @JitendraMalikCV and @pabbeel. Thank you all! 🙏
Code is fully released; check out our:
Homepage: https://t.co/I8kVJO3XPw
Paper link: https://t.co/BFqpOwXYbT
Github: https://t.co/345eTpTx0y
🤖 What if a humanoid robot could make a hamburger from raw ingredients—all the way to your plate?
🔥 Excited to announce ViTacFormer: our new pipeline for next-level dexterous manipulation with active vision + high-resolution touch.
🎯 For the first time ever, we demonstrate ~2.5 minutes of continuous, autonomous control—combining active vision, high-res touch, and high-DoF robot hands SharpaWave — to complete complex, real-world tasks.
Code is fully released; check out our:
Homepage: https://t.co/I8kVJO3XPw
Paper link: https://t.co/BFqpOwXYbT
Github: https://t.co/345eTpTx0y
We then explored the full capabilities of our system—and found it can handle super long-horizon tasks end-to-end.
🍔 Sit back and enjoy the hamburger-making policy in action!