I’ve spent well over 10,000 hours studying math in my life, yet I can’t understand these proofs, at least not without weeks of digging deep into each topic. What’s more, none of my math PhD friends know much about these problems either, and they can’t verify most of them without working directly in the field (yes, math is VERY diverse).
LLMs are getting smarter than the experts themselves, and I’m not sure we have enough bright human minds to verify everything that will come out of them in the coming years.
Remember when we compared AI intelligence to PhD students? I think we’re past that.
If any human mathematician solved one or two of these, that would be notable; solving all 10 would be legendary: one of the all time greats
Now OpenAI is dropping something on this level almost weekly, and (almost) no one pays the slightest attention
Partly this has to do with humans mostly being interested only in other humans (status games); partly it has to do with the reified nature of these problems (most are “pure” mathematicians will essentially no practical implications, at least in the immediate future
But very clear that anything verifiable now belongs (or soon will) to the domain of solved problems
You’ll still have a job, though: disheartening and relieving in equal measure
yes, nonsofic groups exist: this statement is one of many new beautiful results proved by Astra, our next major model.
We're releasing 10 such Astra proofs, complete with lean certificates and CoT walkthroughs for each of them. The results are wide-ranging, from von Neumann algebras (disproof of Connes' Rigidity Conjecture) to better bounds for high dimensional sphere packing, for circuit complexity, for monochromatic triangles in multicolored graphs, and more.
More thoughts here: https://t.co/8SjXONeh38
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
We believe it will be a major step for scientific reasoning. https://t.co/iP6cyheZ7i
HOLY: OpenAI says its *unreleased* Astra model (GPT6?) produced ten advances on long-standing open problems across mathematics, quantum complexity and theoretical computer science.
Among them:
– The first explicit non-sofic group
– Connes’s rigidity conjecture disproved
– Quantum parallel repetition proved for general two-player entangled games
– Ehrhart’s volume conjecture proved
– The first improved general sphere-packing exponent since 1978
OpenAI says the core arguments were generated by Astra. The model then formalized the proofs in Lean, producing machine-checkable certificates alongside a 249-page manuscript.
The successful solution runs would cost only roughly $2,000 in tokens at Sol API rates.
Scientific reasoning is becoming a genuine model capability much much faster than most people expected.
I am so freaking hyped. Breakthroughs every day. The day before yesterday, an 80% price cut for Terra and Luna; yesterday, the DeepSeek 4 flash release with insane evaluations and prices. Today, more breakthroughs with an unreleased model.
I love it! OpenAI is on such a great run!
Foreshadowing from yesterday. Open AI suddenly increasing their stack efficiency and slashing prices. The steadily increasing cadence in model releases. The sudden breakthroughs in math. It's all the same thing. Skeptics, it is time to bite the bullet. We are taking off.
ten significant advances in mathematics and theoretical computer science.
solved using an internal version of Astra, our next major model, for a total cost of about $2000 at Sol API prices:
This would almost certainly be the model that solved the Erdős problem, the same one OpenAI wrote about in 'Safety and alignment in an era of long-horizon models'. I thought they might name this one GPT-6, but what's in a name really. Astra is a great one anyway.
ChatGPT is moving closer to becoming a browser not just an AI that searches the web.
A new update adds URL suggestions while users type, makes it easier to revisit previously opened pages, and provides controls for managing browsing history.
These may look like small navigation improvements, but they fit OpenAI’s larger strategy. The company is retiring its standalone Atlas browser and bringing features such as improved navigation, multiple tabs, downloads, and account logins directly into ChatGPT and Codex.
Instead of convincing people to adopt a separate AI browser, OpenAI is gradually turning ChatGPT into the layer through which they navigate the existing web.
Science moves faster when more researchers have access to the best tools.
Getting frontier AI into the hands of academic researchers means more shots on goal against humanity’s hardest problems — and more breakthroughs that benefit all of us.
Excited to see what they discover.