Boo! 👻 GHOST has been accepted to #RSS2026!
Can we learn manipulation skills from human video (without action retargeting)?
Yes! GHOST learns generalizable manipulation skills by training a hierarchical policy with an embodiment-agnostic sub-goal predictor + embodiment-specific controller.
🧵(1/7)
We will be presenting at CoRL 2025; come see our Oral Session talk (2-3PM) & Poster (4:30-6PM) on 29th September!
Huge thanks to my co-authors @ambermli, Angela, @yu_lifan5260, @BenAEisner, and our amazing advisors Prof. @davheld & Prof. Maxim Likhachev. See you in Seoul! 🇰🇷
Come check out our poster at Poster Session 1 on Sept 28 @ 4:30 - 6:00PM at #CoRL2025!
Website: https://t.co/znEWXekoFo
Paper: https://t.co/oCMic1lHx6
Code: https://t.co/qlqpWdzaj3
Thanks to collaborators: @shhmxy2, @Ying_yyyyyyyy, @ktsim01, @BenAEisner, and @davheld.
Introducing ArticuBot🤖at #RSS2025, in which we learn a single policy for manipulating diverse articulated objects across 3 robot embodiments in different labs, kitchens & lounges, achieved via large-scale simulation and hierarchical imitation learning.
https://t.co/A0SZbAzhAh
🧵
Definitely lots of blood, sweat, and tears went into this one! Congrats to the whole team on getting 2.0 Flash shipped, I'm very grateful to have had the chance to work with everyone on this!
Introducing TAX3D, in which we extend relative-placement methods to generalizable deformable manipulation! #CoRL2024 (1/🧵)
Our approach generalizes to:
- Diverse unseen objects
- Diverse unseen configurations
- Multimodal placements
1/ 🎉 How can we develop methods to generate synthetic, photorealistic data for training #AI models in robotics? We present SplatSim, a step in this direction. SplatSim is a scalable framework that generates photorealistic data for manipulation tasks using existing simulators as a physics backbone — enabling zero-shot Sim2Real policy transfer for RGB policies! 🌍
Check out our project page: https://t.co/JKuHcBWR7h and paper on arXiv https://t.co/p0tKrazgYJ
With @_sparshgarg_ Francisco Yandun, @davheld, George Kantor, @abhi_silwal@CMU_Robotics
@__tinygrad__@realGeorgeHotz Wow, that's wild. The FP16 + FP32 accumulate details are publicly posted for the 4090 (ada whitepaper), but the L40S documentation doesn't break it down into accumulation type. Is that published? Or did you just have to benchmark to figure out empirically what was going on?
@realGeorgeHotz 2. Any insight into why the 8x L40s machines in this benchmark tinybox posted (https://t.co/D17r4483XY) train resnet 2.5x-ish faster than tinybox, even though the total fp16 flops on an L40s node should be like only 50% more? Is it just GPU memory + cpu configuration?
@realGeorgeHotz Super excited about this - considering buying one for our lab's AI+robotics research. Having trouble doing flops/$ calculations, though. 2 (maybe naive?) Qs:
1. Why is advertised NVIDIA tinybox flops 991 tflops, when each 4090 theoretically has 330 tflops of pure fp16 ops?
@sweetgreen seems that there’s something buggy w/ the sweetgreen app this week. 2x I’ve tried to place outpost orders this week, which charged my card via Apple Pay but then immediately refunded saying something went wrong. But my daily Sweetpass credit was still used up. 😢
My friends made a really, really awesome photo-to-Etch-a-Sketch robot. https://t.co/PJu9eerDKH. Draws human sketches at lightning speed - rarely see such high-quality 1x-speed robots. Amazing work!
@asoare159 Yeah I've thought about using diffusion/generative modeling for either the distance directly or diffusing some part of the latent representation, but still need to try it out. Since the points themselves are highly highly correlated, I'd bet that latents would work better.
(1/N) How can we get robots to make precise placement predictions when solving rearrangement tasks, just by watching demonstrations?
In our ICLR 2024 paper, “Deep SE(3)-Equivariant Geometric Reasoning for Precise Placement Tasks”, we do just that!
Paper: https://t.co/GpplhL4Nvf
LLMs are capable of high-level planning, but they require pre-trained skills! Our #ICLR2024 paper instead uses LLM guidance to train RL agents from scratch to solve 25+ long-horizon robotics tasks across four benchmarks w/ >85% success rates
Paper & code: https://t.co/WAP8eNzkzN
@asoare159 Yes it could be used to predict intermediate poses! We've used this to do say, pre-grasp - grasp sequences, similar to your suggestion. The only tricky bit (which we're working on) is handling multimodality natural in most free-space motion seen in demos.
(8/N) We are presenting this work at ICLR 2024! Come swing by our poster at Poster Session 4 on Wednesday @ 4:30-6:30pm.
ArXiv: https://t.co/GpplhL4Nvf
Website: https://t.co/yOWRb2Enfj
Code: https://t.co/Kg5FRhTrkn (release this week)