Wow, I'm honored and humbled that our 2016 perceptual losses paper with @AlexAlahi and @drfeifei was recognized with a Test of Time Award by @eccvconf!
It blows my mind that this is still a core ingredient in many modern pipelines, including the VAEs used in diffusion models
With all the Atlas fun, I almost forgot to tweet about our awesome ECCV tutorial tomorrow. We have great lineup and we are going to cover everything evals in vision and world modeling and robotics.
Wed Sep 9, 1-5pm Malmo time.
https://t.co/Rnv1GsmB4A
Very excited for #ECCV this week! Although I will be virtual, my poster and talk will be there๐
- Sep 9 at 3pm I'm giving a talk on the evaluation of world models at the Evaluation of Evaluating Visual Foundation and Worldย Models tutorial (https://t.co/vIC2RkVERy)
- Sep 11 at 10:30am (poster session 3), @rogerioagjr will present โOut of Sight, Out of Mind? Evaluating State Evolution in Video World Modelsโ (https://t.co/PovQuYu7Y2)
Come check them out!
Todayโs video world models โsimulateโ the world by generating pixel frame observations๐ผ๏ธ. Can they continue to simulate the world when observations are interrupted - such as by occlusion, illumination dimming, or camera lookaway?
To probe this question, we release STEVO-Bench, which holistically evaluates whether image-/text-to-video models and camera-controlled video models can correctly evolve states under observation control. Check out our website, blog and paper for how they fail!
Next view prediction is the key to Atlas, enabling us to unify pixel-level generation and reconstruction. @jcjohnss@BenMildenhall@martin_casado and I had a deeper discussion on some of the most exciting technical innovations of Atlas, our newly released world model for spatial intelligence!
When we say Atlas has pixel-perfect camera control, we mean it.
Atlas precisely follows input camera parameters, including non-planar projections such as the Brown-Conrady distortion model and the Kannala-Brandt fish-eye model.
@BhamidipatiPan1 invented a novel method of camera conditioning and it works beautifully.
๐งต [1/N]
I'm particularly excited by Atlas's ability to reconstruct scenes from a very small number of input image -- was able to create this flythrough of London's Natural History Museum by combining 3 input images I found from completely separate sources on Google Images
Today we sharing Atlas, our new multimodal world model.
One reason this model is so special to me is that it combines two core visual intelligence tasks I have worked on for over a decade: generation and reconstruction
Happy to share what I've been working on since I joined World Labs early this year: Atlas, our next-generation world model for spatial intelligence! Atlas can generate a world, reconstruct a real one, and simulate how it changes through space and time.
The range is hard to describe in text. Stay tuned, I'm gonna share here a bunch of cool gens I made over the last few weeks. ๐บ๏ธ๐คโจ
Two of us filmed this on a couple phones, and we used Atlas to turn it into a short film. This is just a clip from the @theworldlabs tech blog
Atlas is able to generate the parts of the scene that none of the cameras actually filmed.
It can freeze time, allowing you to reframe shots to positions never filmed, then resume.
View events from impossible angles because Atlas knows what was there.
A world model promises one thing: predicting how world operates & changes over space and time.
It's extremely hard to completely solve this (as it pretty much means modeling the whole universe and beyond), but Atlas has made a sizable step toward it.
You know something is right when things used to be impossible start to just work and feel easy. I'm especially impressed by the level of precision it offers. Sparse reconstruction results at this level of completeness and quality were unimaginable before (e.g. https://t.co/9btw4QpiOp)
Can't wait to bring the power of atlas to everyone soon.
Atlas can do all the really hard things with camera-controlled frame gen - orbits, large rotations, getting close to objects without degenerating etc., along with everything you get from an omni ar model. And it can give you a pikachu on the canvas after turning 180๐!So proud to see every part coming together, and will always marvel at what scale (with a lot of care) can do!!
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
When SceniX joined World Labs, we said spatial intelligence was never only about perceiving and generating virtual and physical worlds, but also interacting with them. Today, weโre sharing early results from that vision: building worlds that train robots. ๐๐คโ
The world is not just made of words, and spatial intelligence was never just about perceiving and generating worlds. It's about interacting with them.
Today, SceniX is joining World Labs. ๐๐ค๐
Most inference time scaling approaches are either naive or use fixed recipe for all the prompts. Our #ECCV work combines 3 knobs to get a good generation out fast!!
Paper has lots of analysis on issues with current approaches, how to do inference scaling right and also how to pick a good rater for the generations.
I had lot of fun working on this with @RawalRuchit and other awesome collaborators.
[1/8] Introducing TaskNPoint!
TLDR: A simple, tuning-free recipe for teaching humanoid to perform dynamic skills from a handful of iPhone videos of human demonstrations, trained in under 1 hour on a single GPU.