Next view prediction is the key to Atlas, enabling us to unify pixel-level generation and reconstruction. @jcjohnss@BenMildenhall@martin_casado and I had a deeper discussion on some of the most exciting technical innovations of Atlas, our newly released world model for spatial intelligence!
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence:
LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century.
The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three.
In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete.
00:00 Intro
01:50 The Matrix slow motion scene now takes three iPhones
02:48 Why new view prediction is the primitive
07:10 Unifying generation and reconstruction
11:15 Gaussian splats became the bottleneck
14:17 Dense capture used to mean 300 photos
17:30 Why reconstruction needs generation to fill the gaps
18:44 The LLM lesson image models missed
23:39 The video that made them go all in
28:04 3D design is 95% revisions
30:50 The problem in robotics is data, not chips
32:48 Why a robot policy can't be trained like an image model
34:44 When the simulator becomes the planner
36:45 Frozen time required footage full of movement
40:57 Why new view prediction is AI-complete
42:43 Nature gave animals eyes but not trees
YouTube: https://t.co/AvR59efen0
@drfeifei@jcjohnss@BenMildenhall@theworldlabs@martin_casado
Exciting day for NVIDIA and @huggingface.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI.
Thank you @ClementDelangue for coming to me.
NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗
https://t.co/q8Om2Xc5ye
Valve is celebrating its 30th year anniversary on 24th August. Will that be the release date for the Steam Frame?
(... and can we still dream about a new Half-Life VR game?)
#VirtualReality#VR#SteamFrame
Introducing Bot Mode for Hermes Desktop.
Your agent profiles become a series of named Bots. Each Bot has its own role, model, memory, skills and profile picture; Bots can use any model and even communicate with each other.
Build a specialist Bot once to use it forever.
Today we're introducing Gemini 3.7 Flash, our most intelligent workhorse model yet for coding and agents.
This model brings substantial gains across software engineering, web development, and complex knowledge work.
Now through the end of the year, Gemini 3.7 Flash is available at an introductory price* of $0.75/1M input and $3.75/1M output tokens — making it easier and more cost-effective for developers and customers to scale production-ready agents.
One of the most important updates coming out of Made by Google today is our new sign-to-text feature in Gboard & Live Transcribe.
Built in partnership with the Deaf community, this can help people who sign ASL communicate more seamlessly with their phone or with those who don't sign, like in this Live Transcribe demo:
This soldiering training is the most impressive and immersive I've ever tried. It is 10x better than a YouTube tutorial video. It really allowed me to see the procedure from all the points of view, and even get super close to get the details.
It is a collaboration between @gracia_vr and Imperial College London: they recorded a soldering session with Gaussian Splatting Videos (4DGS), so that you can enjoy it from your VR headset. You can see the action happening in front of you; you can pause and re-watch what you need, change the point of view, get closer, get more distant. And the quality with which it has been shot is impressive: I enjoyed this experience with my DELL Pro Max Tower T2 with NVIDIA Pro RTX 6000, and I could really see all the small details of the PCB that was being soldered!
I was really impressed by it. But I also noticed some drawbacks. First of all, some scenes have artifacts that make seeing the details of the PCB hard. I think when it comes to training involving small details, the capture and reproduction of the splat should be flawless. Then, as much as I loved it as a passive experience, I would have liked to have also some sort of practice session in VR. The power of VR is to let you learn by doing in full safety, and this kind of training does not exploit it. But maybe the best way to enjoy it would be in MR, where you have this training video close to a real workbench where you do the actual soldering while following the tutorial.
Gaussian Splat can really revolutionize training: this kind of recording is much better than any flatscreen experience, and more accurate than any 3D CGI recostruction. I suggest you give it a go at this experience in the Gracia app (it's free). Then let me know your impressions!
#VirtualReality #training #DellProPrecision #GaussianSplat #technology
[Disclaimer: I'm a DELL Pro Precision Ambassador, and this is why I mentioned the model of my PC. I have been given a PC to do cool tests and share my results on social media. No monetary compensation or sales affiliation is part of the collaboration]
Artificial intelligence has turned everyone into a creator. From music to literature, we tracked how the technology is changing different fields. Discover the full story. Register to read for free https://t.co/FbXKJ6lIZb
You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff.
If nothing else, click this link to the 18 minutes in & see how the agents spoke with each other. Its eye opening. https://t.co/G12N4FKfHG
Announcing Discovery Loop!
I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor.
♾
Learn more at: https://t.co/Rv3LMdLluK
This is a port of OpenBrush to #WebXR done entirely using GPT5.5.
Links to everything here 👇
Felix Zhang wrote personally zero lines of code and entrusted his agents to the whole process that took 29 hours.
You can test the app here: https://t.co/LgTnGfYOTh
The full Prompt is here: https://t.co/xq9KdmQXDJ
Here is the GitHub: https://t.co/sghknjVyWd
You truly have no more excuses not to port your Quest app to the web.
#MetaQuest #IMSDK #MetaHorizonPartner
OpenClaw is maturing.
Today we’re introducing monthly extended-stable releases with backported security and reliability fixes, along with a public maturity scorecard for tracking which features are ready for critical workloads.
https://t.co/czpKnKVKCY
Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community.
During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion.
That’s why we created the Open Secure AI Alliance.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb