Introducing Modality Forcing, a recipe for post-training T2I models for SOTA RGB-Depth generation!
Text-to-image (T2I) models learn rich representations of the spatial world.
How do we build on this prior for high-quality depth generation?
https://t.co/uJjGHNiDBu
๐งตย [1/6]
Aux Pays-Bas pour faire avancer lโindรฉpendance technologique europรฉenne. Heureux de retrouver le Premier ministre @MinPres Rob Jetten pour porter ensemble cette ambition.
Today we are introducing Atlas! An autoregressive and multimodal DiT for generating image and depth frames. Atlas is a state-of-the-art 3D reconstruction model and outperforms all open-source models. @HaoZhang623 and I had lots of fun pushing the 3D capability of this model; can't wait to see you build with it!
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
Some of my favorite generations with @theworldlabs Atlas! Really impressed by the model's consistency and ability to handle extreme viewpoints. ๐จโ๐พ๐ฎ๐ค
Another strong demo of Atlas reconstructing the Porsche museum! We have tested this model extensively and are really excited by the performance compared to open models :)
Atlas is out! What excites me most is the 3D capability, months of work with @bardienus : an AR omni model that beats every specialized model at 3D reconstruction. First clip: fly through the Porsche museum, 500+ frames streamed into a single consistent point cloud; feed-forward Gaussian splats follow.
[1/6] Can one model fix rendering artifacts from any 3D representation? And what does it even mean to "fix" a rendering?
At #ECCV2026, we present FixAnything (https://t.co/R4KlUBDt8v): a single generalist video model that refines 3DGS/mesh/sparse point-cloud renderings into photorealistic, 3D-consistent videos.
A huge thank you to everyone who joined @SundaeRobotics 05! ๐ค๐จ Special thanks to @bardienus for an outstanding talk on 3D World Models, Spatial Intelligence & Adaptive AI.
Next Sundae (https://t.co/2HKk7qFV1S), we're excited to host @DanielDugas14 (AI Researcher at Flexion Robotics, previously at Meta FAIR) for V-JEPA 2 & Predicting Physical Intelligence... from self-supervised world models and latent prediction to robot planning, value learning, and Physical AI post-training.
Thanks @PaulKefer@thejonohart@menemazarakis@Joshuabrowder@drjingxi@angelajiazhang@BlaineGame
A huge thank you to everyone who joined @SundaeRobotics 04! ๐ค๐จ Special thanks to @kaiwynd for an outstanding talk on the future of Robotic Simulators.
Next Sundae (https://t.co/n9uixRIjSp), we're excited to host @BDuisterhof, a final-year Ph.D. student at CMU advised by Prof. Jeffrey Ichnowski, for a talk on 3D World Models, Spatial Intelligence & Adaptive AI with Modality Forcing & 3D Pre-training Objectives.
Thanks to @angelajiazhang@thejonohart@menemazarakis@drjingxi@Joshuabrowder
A huge thank you to everyone who joined @SundaeRobotics 04! ๐ค๐จ Special thanks to @kaiwynd for an outstanding talk on the future of Robotic Simulators.
Next Sundae (https://t.co/n9uixRIjSp), we're excited to host @BDuisterhof, a final-year Ph.D. student at CMU advised by Prof. Jeffrey Ichnowski, for a talk on 3D World Models, Spatial Intelligence & Adaptive AI with Modality Forcing & 3D Pre-training Objectives.
Thanks to @angelajiazhang@thejonohart@menemazarakis@drjingxi@Joshuabrowder
Can robot learn manipulation skill with one object and perform it on a completely different one without any new demonstrations?
Excited to introduce SemAnCorr, a framework for dense correspondence and manipulation skill transfer across diverse objects.
https://t.co/p2cLar5vLq
VLAs can do a lot zero-shot. But when one fails, there's no way to tell it what went wrong - you collect more demos and retrain.
What if you could just say something?
Introducing ARCHITECT ๐ค๐ท: a robot coding agent that synthesizes policies from free-form language instructions.
https://t.co/stE4uAdTwT
๐งต1/N
Text2TactileGraphics makes it easier to create tactile graphics for blind and low-vision users! ๐จ๐๐๏ธ
From a text prompt, we generate 3D-printable reliefs with shapes and nuanced textures that anyone can explore by touch.
๐ https://t.co/39zBlqUFGl
๐ https://t.co/vvtT8dx4XR
Recently I've flipped from being bullish to being bearish about AI.
I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so, I would prefer progress) and my reasoning:
The whole 'it turns out if you keep training and scaling the models more they develop broad new capabilities in lots of domains' thesis is wrong (sorry Demis). The recent batch of models haven't got more general, they've got less general. This is most obvious in the fact that their language outputs have got much worse in comparison to e.g. o3. If they were gaining generalist capacities we would expect them to be describing their work in ever more graceful and comprehensive prose!
The image that was being shared as the AGI thesis (November 2025, Tomas Pueyo) was the spiky bubble that has a current spike or two out past human capabilities (e.g. on coding or math) but below human on other capabilities on the other spikes - the future prediction was that as the models scale/advance, every spike would grow bit by bit until the whole center encompasses the human capabilities, with super-superhuman on some spikes. I think it seems like what's actually happened in the last few models has been that the coding/math spike has grown, but leaving behind or even at the cost of the other spikes. The models are no better at some simple logic, language (and sometimes worse!).
This makes sense from a simple RL perspective; you can't RL something endlessly on one domain of tasks and expect it to improve on the other tasks. The fact that early LLMs did seem to improve generally was a byproduct of the written language corpus covering everything - that corpus is general, so training it on that gave the appearance of something generally intelligent and becoming more generally intelligent as it got better at replicating that corpus. But the actual logic and underlying ground truths behind the language aren't captured efficiently enough and weren't effectively RLd in - they top out at some point (I guess this happened around the time that there was the 'has scaling hit a wall' discussion in late 2024). Chain of thought was then a genuine breakthrough, along with web search, which plugged into that general LLM global-corpus intelligence to lead to post 2024 gains.
The AI companies have since worked out that coding works (and pays) really well (basically this is because the entire job is nearly perfectly recorded and exists as training data, and you can set up clear benchmarks and rewards). The recent models (and benchmarks) have been maxxing that and we've seen degradation on normal English use for that reason. This could still be transformative, leading to extremely powerful (and potentially dangerous, particularly in cyber security) models but it's not a pathway to AGI.
I'm probably at about 40% confidence about this. It fits my current observations of AI progress and has a basic explanatory model. It doesn't account for potential breakthroughs, which is a major reason for discounting.
To make some predictions, I guess if I'm right this will become broadly apparent and more widely acknowledged in the next year or two, as we see how the spikiness of models that keep getting released develops.
Maybe there will be efforts to concentrate on specific spikes e.g. health or law which require going back to earlier models and RLing on a different data set/with different rewards/benchmarks. Maybe those separate models can be linked together to give a more apparently general model. How capital intensive that is/the potential profitability will be a defining question. But I just don't see general abilities emerging atm, and I don't think we will any time soon. Good news - a whole industry of tackling important specific problems/sectors can open up!
Introducing FLUX 3.
One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.
FLUX 3 Video is now available in early access (link below).
Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.