Hungry for more Astra robot videos? 🤖
How about a large-scale sandboxed evaluation to go with it? We ran 98,000 evaluations across 28 simulated environments and found some surprisingly clever physical reasoning along the way!
New preprint 🧵👇 (1/9)
Hungry for more Astra robot videos? 🤖
How about a large-scale sandboxed evaluation to go with it? We ran 98,000 evaluations across 28 simulated environments and found some surprisingly clever physical reasoning along the way!
New preprint 🧵👇 (1/9)
Hungry for more Astra robot videos? 🤖
How about a large-scale sandboxed evaluation to go with it? We ran 98,000 evaluations across 28 simulated environments and found some surprisingly clever physical reasoning along the way!
New preprint 🧵👇 (1/9)
@chrisdotai It's also worth noting that agents still have interactive access to the simulator, so they can use that to generate variants and search for edge-cases. For example, in the paper we show how in this environment agents explicitly look for hard seeds and use them for testing.
@rb_rupam We mainly used the environments from the KinDER benchmark (https://t.co/i3W5zH3qw3), they are mostly implemented in Mujoco/PyBullet, so no generative models involved on that side
Over the last six months, coding agents have changed how I do research. This project has made me realize how much we need to understand them, beyond using them as tools.
It's also helped me understand my own research taste and find a direction I'm really excited about. (8/9)🧵
@basisorg and my group at Princeton are recruiting a postdoc to work at the intersection of robotics and code-based world models.
We’re looking for someone excited about abstractions & planning, and who knows their way around a real robot.
Link 👇 Thanks for boosting!
New blog post:
https://t.co/T29LEH7lpY
This is the second in a "series" of posts about how to define and work on research problems. (The first post was 3 years ago...)
This one addresses the agentic elephant in the room.
🤖 Curious about the connection between Task and Motion Planning (TAMP) and robot learning?
📄 Check out our new survey on Learning By and For Task and Motion Planning!
We organize the growing literature on TAMP + learning, introduce a taxonomy of TAMP learning, clarify the state of the art, and highlight promising directions for future research. 🚀
🔗 https://t.co/uqha9QP4gf
Long-horizon tasks + sparse rewards = no supervision. So, we need dense supervision methods… but how do we evaluate them?
Meet QVal, our latest release👇
Meet KinDER — a stress test for robot physical reasoning. All 13 methods failed 😈
🌎 25 environments
♾️ Infinite tasks
🏋️ Gymnasium API
⚒️ Over 20 parameterized skills
🪧 Human demonstrations
📊 13 baselines (planning and learning)
From @Princeton@CMU_Robotics@ICatGT@CambridgeMLG@nvidia@MIT_CSAIL
🧵 1/n
I'm happy to announce our paper
Generating Code World Models with Large Language Models Guided by Monte Carlo Tree Search, has been accepted @NeurIPSConf'24!
Link: https://t.co/R8Si9oohje
Joint work with @dainese_nicola, @MerloMerler and @marttinen_pekka.
#NeurIPS2024