When we say Atlas has pixel-perfect camera control, we mean it.
Atlas precisely follows input camera parameters, including non-planar projections such as the Brown-Conrady distortion model and the Kannala-Brandt fish-eye model.
@BhamidipatiPan1 invented a novel method of camera conditioning and it works beautifully.
🧵 [1/N]
We cooked a new model at @theworldlabs! Atlas, our next-generation world model, is out today and I am so excited to finally talk about this one! 🔥🌍
The main idea and unique feature of Atlas is camera-conditioned sparse frame generation: give it a one or more images and the exact camera poses you want, and it generates frames at those poses, consistent in 3D with everything it has seen. You can do also many things with this model, like
(1) Sparse reconstruction: even 2-3 photos of an apartment or warehouse are enough for making a "digital twin".
(2) Keyframing for film-making: place a handful of frames in 3D space and direct a 1 minute video at 1440p through them. You can frame every shot yourself instead of pulling the lever of a slot machine.
(3) Bullet time (my favorite!): You can freeze the time mid confetti explosion and fly the camera through it, from footage of 3-5 cell phones on tripods. The Matrix needed a warehouse of cameras for that shot, now it fits in a backpack.
Because every task is just a sequence of frames + cameras, we trained Atlas on a big mix of tasks and data. I personally drew a lot of satisfaction in pouring SO MUCH DATA into this big boy. Due to task generality, Atlas can also do text-to-image, 360 panoramic generation, outpainting, depth estimation, etc. Only my teammates know how many bernese mountain dogs I have generated with this model in text-to-image mode. XD
A huge shoutout to the entire team at World Labs, I feel so lucky to work with such talented people and have made long-lasting friendships being in the trenches together, day in and day out - @KeunhongP@_mbanani@BhamidipatiPan1@_nikunj_gupta@gowthami_s@rmgarg02@vibhaa14@wendlerch@hyl0m0rphism@DrTunnels@bardienus and many more (not all are on X).
Blog: https://t.co/K1UUtlTB4Z
See thread below for details 🧵⬇️
Atlas has one operation: predict what a camera at a given pose would see.
Novel-view synthesis, panorama generation, in-painting, out-painting, super-resolution, text-to-image: the field treats these as different problems. Atlas treats them as the same. The camera tells the model which pixels to generate:
— point the camera somewhere new → novel view synthesis
— faces of a cube → panorama generation
— withheld pixels inside the frame → in-painting
— pixels past its edge → out-painting
— the same view, denser → super-resolution
— no input image → text-to-image
Atlas is live! A spatial foundation model, trained in-house from scratch.
After shipping RTFM last year, I was convinced a single camera-conditioned model could unify most spatial generation tasks. Atlas is that bet at scale.
The core idea👇
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
World Labs has raised $1 billion in new funding. We are grateful and excited to partner with our investors, including AMD, Autodesk, Emerson Collective, Fidelity Management & Research Company, NVIDIA, and Sea, among others.
https://t.co/MfK7mLGKPH
@ericsmith1302 Congrats @ericsmith1302 , does it also have an option support AI generated clips? Also what is the best marketing channel that worked for you?