World Labs CEO @drfeifei on the need for world models:
"There's no language out there. You don't go out in nature and there's words written in the sky for you."
"There is a 3D world out there that follows laws of physics, that has its own structures due to materials. To fundamentally back that information out and be able to represent it and be able to generate it is just fundamentally quite a different problem."
"We will be borrowing useful ideas from language and LLMs, but this is fundamentally, philosophically a different problem."
Latest Deep RL class lectures are now online!
https://t.co/GvqI1v3hgD
Thanks to @seohong_park, we now have CS185/285 for spring 2026 available to everyone to watch.
Course website here: https://t.co/U16rasTOAo
Apologies for a few recording glitches (it's not a perfect system).
Now in preview: The ChatGPT desktop app for Linux.
Use ChatGPT, ChatGPT Work, and Codex where you already work and build, with your projects and browser workflows on supported Linux systems.
Ended neurips with this score 😭, how do one even get these scores , 4 ,5 and 1 (not even a 2 or 3 )....
I dont think this will be accepted so gonna take the feedback during reviews and submit it elsewhere .
An internal version of our next major model produced 10 new results on long-standing open problems in mathematics and theoretical computer science, using roughly $2,000 worth of tokens at GPT-5.6 Sol API rates.
An internal version of our next major model produced 10 new results on long-standing open problems in mathematics and theoretical computer science, using roughly $2,000 worth of tokens at GPT-5.6 Sol API rates.
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all.
I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand.
Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta!
🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇
🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!
Check out the configuration details in our official API docs: https://t.co/smCwQZMeiq
.@Alexandr_Wang's advice to his 18-year-old self: develop your own internal compass for how the future will unfold, and hold conviction in it against the noise.
At Startup School 2026, the Scale AI (YC S16) founder — now leading @Meta’s Superintelligence Labs — talks with @garrytan about rebuilding a frontier lab from scratch, why talent density compounds, and how to spot the exponential worth betting your twenties on.
00:07 — How Alexandr Wang Started Scale AI
03:25 — Pivoting to the Right Idea
06:23 — Conviction Before Consensus
09:06 — Why This Is the Best Time to Start a Company
11:27 — What Personal Superintelligence Looks Like
13:10 — Building a Frontier AI Lab
16:36 — Why AI Models Need to Be Cheap
20:01 — Vision Will Matter More Than Intelligence
24:06 — Systems Thinking in the AI Era
26:51 — The Biggest Opportunity in AI Today
29:25 — Advice to My 18-Year-Old Self
major price cuts today:
*80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output
*20% drop for GPT-5.6 Terra, to $2/$12
*GPT-5.6 Sol gets Fast mode in the API, up to 2.5x the speed for 2x the price, same intelligence
For decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand.
Today, we take a major stride toward making that dream a reality:
Introducing Gemini Robotics 2 from @GoogleDeepMind, the intelligence layer powering the next generation of truly adaptable robots. This major advance unlocks intelligent whole-body control, advanced dexterity, and even multi-robot collaboration 🤯.
Ok but... how does a robot actually "think"?
Real-world tasks take time and planning. To manage that complexity, our new embodied reasoning model, Gemini Robotics ER 2, acts as the robot’s high-level brain, enhancing the robot’s capabilities to:
— Observe the environment
— Reason about the actions needed to complete the task
— Coordinate with the vision-language-action model to carry out actions
— Track progress until the job is done
This setup allows robots to execute complex multi-step workflows, self-correct if a step fails, and adapt to completely novel situations.
Learn more about Gemini Robotics ER 2 (and our two other brand new models) here: https://t.co/1YEpoYAhww