Imagine what we can build when one model can index space and time to generate 3D-consistent worlds! After months of grinding with this incredible team @theworldlabs, it’s so exciting to see Atlas come together!
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
Can point flow completion serve as a pre-training objective for learning to understand physics?
Excited to share PointZero, our recent work on bringing this idea to reality.
My favorite part is that the model works zero-shot on real-world images! Take a photo, and then start to simulate it with control points. An embodiment-agnostic, scalable approach.
Huge congrats to @bardienus and the team!
Today’s video world models “simulate” the world by generating pixel frame observations🖼️. Can they continue to simulate the world when observations are interrupted - such as by occlusion, illumination dimming, or camera lookaway?
To probe this question, we release STEVO-Bench, which holistically evaluates whether image-/text-to-video models and camera-controlled video models can correctly evolve states under observation control. Check out our website, blog and paper for how they fail!
It feels amazing to see the bullet-time results come together, especially when we were all so close to giving up. Huge thanks to @jcjohnss and @BenMildenhall for trusting me!
Next view prediction is the key to Atlas, enabling us to unify pixel-level generation and reconstruction. @jcjohnss@BenMildenhall@martin_casado and I had a deeper discussion on some of the most exciting technical innovations of Atlas, our newly released world model for spatial intelligence!
When we say Atlas has pixel-perfect camera control, we mean it.
Atlas precisely follows input camera parameters, including non-planar projections such as the Brown-Conrady distortion model and the Kannala-Brandt fish-eye model.
@BhamidipatiPan1 invented a novel method of camera conditioning and it works beautifully.
🧵 [1/N]
@sheldonblewis interned with us this summer and made pivotal contributions to Atlas.
If you're looking for someone with superb execution and research sense you should reach out to him.
I’ve learned so much from working with @KeunhongP . He’s a true technical leader and a detail-oriented thinker who excels in technical leadership, implementation, and execution.
Imagine what we can build when one model can index space and time to generate 3D-consistent worlds! After months of grinding with this incredible team @theworldlabs, it’s so exciting to see Atlas come together!
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
Atlas also supports native 3D outputs, beating all specialized models in sparse reconstruction.
Big shoutout to our super interns @HaoZhang623@bardienus @DrTunnels!
🔥An incredible accomplishment. Think of it as a video model with full camera control. And the scene remains (nearly) 3D consistent. Built on a fully internal base model. There are many use cases, from video editing, to 3D reconstruction to robotics. Great technical blog too.