After we have determined all the proofs, discovered all the drugs, found all the new materials, we will have to face the reality that our deepest problems are political, philosophical and psychological.
It turns out DIGGER is a work of pure cinematic genius. Strongly recommend, 4+ stars. Donβt read anything about it or pay attention to anything anyone says, just go.
Look at how much further 30 minutes by flying car gets you than driving.
SF
00:00 β car
00:03 β 120 mph flying car
00:06 β 250 mph flying car
LA
00:18 β car
00:21 β 120 mph
00:25 β 250 mph
NYC
00:36 β car
00:39 β 120 mph
00:43 β 250 mph
OpenAI's internal model clearly has "research taste" in math - i.e., the final component we need to get to RSI.
- jboggan (post on Hacker News) on Barnette's Conjecture: "the 'aha' insight for this is actually f**ing wild... this is the first time I've seen complex roots and annihilating terms like this... I don't understand where this trick originated."
- Joshua Zelinsky on the proof that the chromatic number of the plane is at least six: "doesn't look like the method is a direction that the prior lit used to my knowledge... far beyond merelt building on existing methods or seeing connections between different problems."
and on two other problems (where he says he is only partly familiar with the literature): "not remotely low-hanging fruit... it seems like the AI is somehow inventing new techniques on its own."
These match Tristian Buckmaster's view on the Navier-Stokes solution: "you combine... ideas of convex integration with the growth mechanism of the Euler blowup, and... you create a new mechanism which is used to correct this non-solution. This is actually a cool idea. It's the kind of idea that I've been trying and failing to realize for over ten years... I didn't manage to do it... This is the leap."
The average amount of compute used to solve these problems was ~3 hours of Pro-level thinking.
And so, this means that OpenAI's internal model is able to generate truly novel discoveries in mathematics for... maybe at most a few hundred bucks?
If I were OpenAI, I would be asking this model to immediately target major unsolved problems in AI R&D.
How fast is AI's research taste improving?
We find that the experimental research taste of frontier models has doubled every ~3 months since December 2025. The best model, Opus 5.5, now exceeds our expert human baseline. Our human experts are experienced researchers, but most havenβt worked at a frontier lab.
Why measure research taste? In the AI Futures Model, it largely determines how quickly artificial superintelligence is reached once coding is fully automated.
@andrewparker@noahrshinn Once it works well for them I don't think people will want to switch, especially not if it feels like you have to teach someone new to know you. Also there are network effects here: what Instinct learns for one user it learns for all users.