Real-world capture is almost never complete. A missed corner, object, a room that didn't get enough viewpoints etc. leaving empty holes where the geometry should be.
Our SOTA model Echo is built for that reality. It fills unscanned regions from surrounding visual context, so you still get a complete, explorable 3D world instead of broken voids.
Incomplete capture shouldn't mean an incomplete experience.
@gabriberton My question is what does RL learn beyond a strong handcrafted procedure? We can use tools for dg detection, generate crop/zoom candidates ... Given the data budget required for RL, is it actually worth it here?
@gabriberton It seems the space of useful actions an agent can take here is quite limited. Can RL actually learn something beyond a well-designed set of handcrafted rules and heuristics?
@nanliuuu A lot of physics is latent, not visual. Friction, mass, contact forces don't show up cleanly in pixels. That's where enforced biases can help. Ideally not a permanent substitute, but to compensate for the missing supervision, then shed as the model learns to infer it on its own.
We wired 3D Gaussian Splatting into the 1999 Quake 3 engine - it's fully playable!
Echo-2 generates the game levels, and it's rendering spaces directly inside the open source version of id Tech 3.
Anyone wants to try it?
Image models gave us the power to see our imaginations in 2D. At @SpAItial_AI , we're unlocking the third dimension. I typed a prompt into Echo-2. Five minutes later, I was walking through it. No manual work. Just words → worlds. Try it yourself: https://t.co/pxOIbL3Fq0
🚀Echo-2 is here - our new world model!
These aren’t videos. These are 𝟑𝐃 𝐬𝐜𝐞𝐧𝐞𝐬. Generated from a single image.
- Stunning visual quality.
- Real-time rendering.
- Interactive camera control.
- Physically grounded.
🧵More details👇
Can RL with outcome rewards alone just efficiently explore outside the support of the base model and learn completely new capabilities?
We study it in a theoretically tractable setup, and prove that outcome rewards are not enough, but process rewards can get you there! [1/3]
The concept of creating an exact digital replica of the physical world has always fascinated me: environments that look and behave exactly like our everyday reality, precisely captured in the digital domain. This is the essence of 𝐖𝐨𝐫𝐥𝐝 𝐌𝐨𝐝𝐞𝐥𝐬, simulated realities indistinguishable from our own. Generating these models is the core mission behind what we are building at @SpAItial_AI.
True World Models must capture both photorealistic appearance and underlying physics, spatially-consistent across the environment. For static scenes, current models already deliver impressive results, unlocking downstream applications from gaming to 3D design. However, the true frontier lies in modeling dynamics, which will enable the training of AI agents whose learned behaviors can bridge the sim-to-real gap, thus unlocking countless real-world applications.
🚀 Announcing Echo — our new frontier model for 3D world generation.
Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely.
This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations.
Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment.
Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted.
Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful.
Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world.
This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right.
Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍
#Echo #WorldModels #SpatialAI #3DFoundationModels
Check it out: https://t.co/QgsuLdhoe6
🚀🚀🚀Announcing our $13M funding round to build the next generation of AI: 𝐒𝐩𝐚𝐭𝐢𝐚𝐥 𝐅𝐨𝐮𝐧𝐝𝐚𝐭𝐢𝐨𝐧 𝐌𝐨𝐝𝐞𝐥𝐬 that can generate entire 3D environments anchored in space & time. 🚀🚀🚀
Interested? Join our world-class team:
🌍 https://t.co/U0JNkNwp3s
#GenAI#3DAI
After great years as a PhD & postdoc at MPI Informatics, I’m excited to share what’s next!
I am very honored to receive the Eurographics PhD Award — special thanks to my advisors Karol Myszkowski & Tobias Ritschel (UCL), my collaborators, and my amazing wife & family.
After two incredible years at Meta, it’s time for a new adventure!
It’s been a journey full of growth, fun, and pushing the limits of 3D scene reconstruction & synthesis. I'm deeply grateful to have worked with such a brilliant, dedicated team and for all the wonderful memories.