Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects.
Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below๐งต
We're releasing MolmoMotion, a 3D motion forecasting model.
Given one or a few video frames, 3D points on an object, & an instruction like "Put the white bowl on the table," MolmoMotion predicts where those points will go over the next few seconds in a shared 3D world frame. ๐งต
We believe in the representation of 3d point trajectory as an alternative form of world modeling. it is universal to all object, more compact than pixel, and can directly be consumed by many downstream task.
p.s. the key for us to generate high quality real world data is we respect objectness and do hard work in object grounding -- which allows us to properly filter and smooth noisy trajectories coming out of off the shell 3d tracking methods.
And that's a wrap on a fantastic ICRA 2026! ๐ Incredible run for MolmoBot โ clean sweep on workshops, winning Best Paper at all three we entered: Synthetic Data for Robot Learning, Beyond Teleoperation, and VLA Pipelines. ๐ค
Our paper MolmoB0T won top honor at the SDRL workshop today! Congrats to my co-authors @ab_deshpande, Maya, Snehal, @RanjayKrishna, @shahdhruv_ (plus others, you know how it goes), and we'll see you as well at the VLA Pipelines and Beyond Teleoperation workshops on Friday ๐
MolmoBot, our open robotic manipulation suite trained entirely in simulation, now has code, training data, a data generation pipeline, & evals all available.
This puts our robotics models within reach of any research labโno extensive real-world data collection required. ๐งต
Jensen approves!
Hercules efforts from @YejinKim4@omarrayyann, Max Argus & team! This has a decent chance of becoming a super important benchmark fo robotics going forward.
Check out this @RoboPapers episode with the MolmoSpaces folks.
Today, a step forward in open robotics - our results show that sim-to-real zero shot transfer for manipulation is possible. MolmoBot is our open model suite for robotics, trained entirely in simulation on MolmoSpaces.๐งต
Are you sure your training data is actually synced?
Egocentric camera sees a hand grasping an orange, but the wrist cam shows nothing and tactile reads zero contact.
Your policy is learning from broken data and doesn't even know it.
In Physical AI, multi-modal sync is everything.
โ Egocentric: 30fps
โ Wrist: 30fps
โ Tactile: 100Hz
Different devices, different clocks, slightly different rates. The drift starts small. Barely noticeable frame by frame. But over a 4-minute episode, that tiny difference compounds into seconds of misalignment. And you had no way to even check.
Until now.
We built the Sync Quality Dashboard. One score tells you if your data is clean. Then go deeper. Clock offset, drift rate, jitter, frame drops, per-stream correction. All visible, all measurable.
In a 4-min episode, accumulated clock drift reached 7.5 seconds by the end of the recording. After correction: 9.0ms. That's the difference between "roughly aligned" and "actually aligned."
Visually confirm vision-to-vision, vision-to-tactile alignment frame by frame. No more "trust me, the data is fine."
We don't just collect multi-modal demos. We ship a quality assurance layer so you can verify every episode before it touches your model.
All data in @LeRobotHF format. Ready to train. Verified in sync.
Stop guessing. Start verifying.
MolmoSpaces-Bench leaderboard is now live! Test your generalist policies to see how they compare across tasks and environments. Feel free to reach out if you need help setting it up.
https://t.co/d7RlL6h7ZE
๐๐ซ๐๐๐ฆ๐๐๐ซ๐จ ๐ข๐ฌ #๐ ๐จ๐ง ๐๐จ๐ญ๐ก ๐๐จ๐ฅ๐ฆ๐จ๐๐ฉ๐๐๐๐ฌ ๐๐ง๐ ๐๐จ๐๐จ๐๐ซ๐๐ง๐ ๐
๐ช๐ต๐ฎ๐ ๐บ๐ฎ๐ธ๐ฒ๐ ๐๐ต๐ถ๐ ๐ป๐ผ๐๐ฎ๐ฏ๐น๐ฒ: DreamZero-DROID is trained ๐๐๐๐ ๐ ๐๐๐๐ก๐โ using only the DROID dataset. No pretraining on large-scale robot data, unlike competing VLAs. This demonstrates the strength of video-model backbones for generalist robot policies (VAMs/WAMs).
More broadly, training ๐๐๐๐ฆ on real data and evaluating on (1) transparent, distributed benchmarks like ๐๐จ๐๐จ๐๐ซ๐๐ง๐ or (2) scalable sim-benchmarks like ๐๐จ๐ฅ๐ฆ๐จ๐๐ฉ๐๐๐๐ฌ is an exciting step toward fairer and more reproducible evaluation of generalist policies, one that the community can hillclimb together to measure progress.
Special thanks to the Ai2 MolmoSpaces team (@notmahi @omarrayyann @YejinKim4 Max Argus) and the RoboArena team (@pranav_atreya) for helping with the set-up and getting these evaluations! Special shout out to @youliangtan@NadunRanawakaA@chuning_zhu, who led these efforts from the GEAR side :)
+ We also release our DreamZero-AgiBot checkpoint & post-training code to enable very efficient few-shot adaptation. Post-train on just ~30 minutes of play data for your specific robot, and see the robot do basic language following and pick-and-place ๐ค(See YAM experiments in our paper for more detail).
++ We also provide the entire codebase & preprocessed dataset to replicate the DreamZero-DROID checkpoint.
๐ https://t.co/pEoQ9QrTHW
๐ป https://t.co/AdHqcLuwIy
RoboArena: https://t.co/NiUQxLLL6F
MolmoSpaces: https://t.co/ZXiguGYPx9
MolmoSpaces also comes with 42M+ grasps that cover 48K+ objects across 250K+ scenes, allowing large-scale functional trajectory generation in MuJoCo and IsaacSim.
MolmoSpaces provides singular scale and diversity. We built a benchmark that puts that scale to use.
MolmoSpaces-Bench evaluates zero-shot policies across thousands of environments previously unseen to them under systematic variation, providing insights that go beyond a success rate %
More Below:
Introducing MolmoSpaces, a large-scale, fully open platform + benchmark for embodied AI research. ๐ค
230k+ indoor scenes, 130k+ object models, & 42M annotated robotic graspsโall in one ecosystem.
Excited to present โClimate-sensitive Urban Planning through Optimization of Tree Placementsโ w/ Ferdinand Briegel, Max Argus, @envmet & @ThomasBrox at the #NeurIPS2023@ClimateChangeAI workshop. Tl;dr: we optimize urban tree locations for improved outdoor human thermal comfort.