World models shouldn’t need to search through 9,000 action sequences before they can act.
Introducing INTACT-JEPA: a world model that maps latent intent directly to action.
⚡ 0 search
🎯 95.3% success
🚀 2.9–5.5 ms planning — ~300× faster
World models that don’t just predict the future.
They know what to do next.
🌐 https://t.co/EIvLVULdgm
💻 https://t.co/mcocYN8OFK
📄 https://t.co/k0qr3qC86u
🔁 Gravity 4D closes the loop from 4D imagination to robot action.
🤖 Rather than using 3D/4D features only as perception inputs or auxiliary supervision, Gravity 4D jointly predicts future RGB video and 4D representations—pointmaps for spatial structure and scene flow for 3D motion—and routes the generated 4D latent back into action generation.
The model first imagines how the future world should look and move in 3D, then uses inverse dynamics to produce the corresponding action chunk.
📈 Trained on LIBERO and evaluated zero-shot on LIBERO-Plus, Gravity 4D improves success from 73.73% to 78.62%, with +11.21 pp under camera-viewpoint shifts and +8.69 pp under sensor noise.
🧪 Ablations show that RGB and 4D are complementary: RGB retains semantic and appearance priors, while 4D provides geometry and motion-aware robustness.
📄 Paper coming soon.
🤖 What if every robot could have its own face?
Thrilled that our paper “Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots” is an #RSS2026 Best Paper Finalist! 🏆
Give us one 2D portrait: Yoda, Jack, Rose, an elf, or a legendary Chinese character, and our system automatically:
⚙️ designs a personalized, collision-free facial mechanism
🖨️ turns it into a manufacturable robotic head
🗣️ generates real-time speaking and listening expressions
⏱️ cuts expert mechanical design from 22.8 hours to 11.7 minutes
From a picture to a physical, conversational robotic face.
🌐 Project: https://t.co/VAtBCPwY9k
💻 Code: https://t.co/tvbMuB5rFt
The future may not be one humanoid face copied thousands of times, but thousands of robots, each with its own identity and personality.