@ylecun@haider1 working with spatiotemporal data I feel this in my bones, our models crush static benchmarks yet lose track of a moving scene within minutes, the missing piece is not more text, it is a memory of how the world keeps changing
That video is 2 years old.
Two years later, we are still far from human-level AI.
Sure, AI has superhuman performance in a number of tasks (particularly in mathematics and coding and answering questions with known answers).
But where is my Level-5 self-driving car?
(Before you ask, Tesla FSD is rated Level 2, Waymo is Level 4).
Where is the self-driving car that can teach itself to drive in 20 hours of practice like any 17 year old?
(Before you ask, we have billions of hours of training data and still can't successfully train a self-driving system by imitation learning)
Where is my domestic robot?
Where is the robot that can do what any 8 year-old child can do?
Where is the robot that can learn a new task as quickly as an 8 year old?
AI still has a hard time with the complexity and messiness of the physical world.
The critical distinction between base LLMs (2024 and earlier) and modern LRMs is not symbolic tool use. It's the switch from a transductive paradigm (intuit the answer to the query) to an inductive paradigm (intuit the program/instructions that produce the answer to the query).
They're trained to be inductive, and they perform test-time induction, i.e. test-time prediction of a NL program / reasoning chain. This unlocks entirely new capabilities -- in particular fluid intelligence. Base LLMs, to this day, have ~0 fluid intelligence. LRMs have substantial levels of fluid intelligence.
The performance of LLMs on ARC 1 (a benchmark from 2019) remains ~10-15% today. Scaling them up by a factor ~100,000x got them from 0% to 10%. Meanwhile LRMs the same size or smaller saturated ARC 1 in 2025.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
@ValsAI the held-out part matters more than the leaderboard, a static public set basically becomes training data the moment it ships, this is the first search eval I have seen treat that as the main problem
The idea of an intelligence explosion caused by recursive self improvement has been around for a long time but until very recently it did not seem imminent. Now many leading researchers think it may happen quite soon. You can read our paper about it here:
https://t.co/sgUpugjpRY