Introducing "3D HAMSTER", accepted to IROS 2026! 🎉
Hierarchical VLA planners draw waypoints in 2D, but robots act in 3D. Give the VLM a depth encoder, and it predicts metric 3D trajectories robots can execute.
Project page: https://t.co/5TiGDOmHLB
Paper: https://t.co/Rq3rIkUAP4
Introducing "See like a Robot"🤖
Robot data spans diverse camera viewpoints, making learning harder. Give a VLA robot-centric pointmaps, and it performs better with one extra encoder + one element-wise addition.
Project page: https://t.co/k5uBJOUBYa
Paper: https://t.co/ff2pwLLxxl