Introducing RPDiff - a diffusion model for 6-DoF object rearrangement in 3D scenes
Website (videos+code): https://t.co/J0M0S9rVFi
RPDiff can perform tasks like book shelving, can stacking, and mug hanging, gracefully handling multi-modality and generalizing to diverse scenes 🧵
Love imitation learning but need more robustness? Check out RialTo -- a pipeline for acquiring reactive visuomotor policies via real-world BC and simulation-based RL with digital twin 3D assets
https://t.co/6DqWvQGflh
Project lead by @marceltornev, details in his thread below!
How can we train robust policies with minimal human effort?🤖 We propose RialTo, a system that robustifies imitation learning policies from 15 real-world demonstrations using on-the-fly reconstructed simulations of the real world. (1/9)🧵
Project website: https://t.co/ZOU9zw7rVW
Our new robotic assembly planner “ASAP” has been accepted by #ICRA2024! We propose an automated method to generate physically stable assembly sequences executable by robot arms for general-shaped assemblies, one step towards real-world autonomous assembly: https://t.co/EQRygL7sDX
Shelving, Stacking, Hanging: Relational Pose Diffusion for Multi-modal Rearrangement
paper page: https://t.co/y0zA124gN0
propose a system for rearranging objects in a scene to achieve a desired object-scene placing relationship, such as a book inserted in an open slot of a bookshelf. The pipeline generalizes to novel geometries, poses, and layouts of both scenes and objects, and is trained from demonstrations to operate directly on 3D point clouds. Our system overcomes challenges associated with the existence of many geometrically-similar rearrangement solutions for a given scene. By leveraging an iterative pose de-noising training procedure, we can fit multi-modal demonstration data and produce multi-modal outputs while remaining precise and accurate. We also show the advantages of conditioning on relevant local geometric features while ignoring irrelevant global structure that harms both generalization and precision. We demonstrate our approach on three distinct rearrangement tasks that require handling multi-modality and generalization over object shape and pose in both simulation and the real world.
For more details, check out our paper! “Shelving, Stacking, Hanging: Relational Pose Diffusion for Multi-modal Rearrangement” - https://t.co/2m54EYxje1
For code and results videos, check out the website - https://t.co/J0M0S9rVFi
Introducing RPDiff - a diffusion model for 6-DoF object rearrangement in 3D scenes
Website (videos+code): https://t.co/J0M0S9rVFi
RPDiff can perform tasks like book shelving, can stacking, and mug hanging, gracefully handling multi-modality and generalizing to diverse scenes 🧵
There are many other cool rearrangement works (NSM https://t.co/txovcCrtPY, TAX-Pose https://t.co/p7v3EUuOsD, our own R-NDF https://t.co/3rGHs1Bdwp), that mainly focus on individual objects and/or single-solution tasks, whereas we consider scenes that offer multiple solutions
Love this and its connection to motion planning. Planning is so powerful in its generalization across tasks, but linking perception to "plannable" representations (and having *this* link generalize) is still hard. Here's one very compelling way to do it. Great job @wenlong_huang
How to harness foundation models for *generalization in the wild* in robot manipulation?
Introducing VoxPoser: use LLM+VLM to label affordances and constraints directly in 3D perceptual space for zero-shot robot manipulation in the real world!
🌐 https://t.co/FhBazTzi7Z
🧵👇
Can we figure out where our robot is from touch? MidasTouch, to be presented as a @corl_conf oral, uses vision-based touch to disambiguate the pose distribution b/w a robot finger and known object.
https://t.co/oqJ50zFMxw (1/8)