@an_interstice (In response to Kontsevich suggesting AI soon)
Milner: So we’ll have a problem with the prize? I’m thinking about splitting it between hundreds of people
[everyone laughs] What makes you so optimistic?
Kontsevich: No it's actually pessimistic
[more laughter]
2015!
@an_interstice (In response to Kontsevich suggesting AI soon)
Milner: So we’ll have a problem with the prize? I’m thinking about splitting it between hundreds of people
[everyone laughs] What makes you so optimistic?
Kontsevich: No it's actually pessimistic
[more laughter]
2015!
@lukeisamazing The guy whose work was allegedly stolen described his own work as AI slop! His second author was only on it because of his internal Claude access! Everybody in this story was using substantially using AI!
@MattZeitlin I suspect the lack of engagement is part of the appeal, Terry’s blog is extremely high quality and he just wants a place to fire off more casual posts without morons bothering him. Baez(like many academics)jumped ship post-Elon pre-Bluesky. Others just like the smaller community
@RadishHarmers “we collectively had no research-level expertise in fluid dynamics and the Navier-Stokes problem, and therefore were unable to meaningfully contribute to the mathematical content.”
Seems like they had no one and just trusted the Lean formalization
Some more technical points:
(a) We began working on the Millennium problems due to viral twitter rumors that Anthropic had resolved 2 Millenium problems. Our aim was to see whether our system was also capable of this impressive feat, especially given our excitement regarding the large recent capability increases of our internal model detailed in our blog post.
(b) We did not see any of their work until they released it publicly last night. One can in hindsight see that our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced). It is also worth noting that our team consisted largely of mathematicians, physicists and computer scientists with the deepest roots in academia and respect for the stature and integrity of the Millennium problems.
(c) Levent reached out unprompted to an OpenAI employee on Wednesday. Due to his numerous posts on Twitter (S^6, Jacobian, Hadamard…) which appear to represent Anthropic and use Anthropic internal models, we believed that their project was, at least in part, an Anthropic project. Additionally recent further tweets by Levent (“augustus mirabilis”) led us to believe that they had solved at least 1 Millennium problem and would likely release it soon.
(d) Regarding the level of human involvement on our end: although a group of people was involved in our efforts, we collectively had no research-level expertise in fluid dynamics and the Navier-Stokes problem, and therefore were unable to meaningfully contribute to the mathematical content. We involved several people in order to initially attempt multiple Millenium problems using a variety of multiagent techniques.
@JohnBal14046363 $500k of o3 ~= $20 of Astra twenty months later, prices for *these capabilities* will fall. Ofc, we are obviously still a long ways from the far tail of capabilities, and it seems clear the frontier will be perpetually dominated by those with the $$$ for it in no time at all
Yes, this result cost millions of dollars.
But remember that when @OpenAI announced o3 it cost ~$500,000 to score 87.5% on ARC-AGI 1. Today, Astra scores higher for ~$20.
In 2025 it took us and GDM an enormous amount of compute to achieve IMO gold. For the 2026 IMO, anyone with a $20/month ChatGPT subscription could do it.
Massively scaling test-time compute gives us a glimpse of the future. I believe that a year from now everyone will have an AI at their fingertips capable of solving problems of this caliber.
If I am reading this correctly, Tristan Buckmaster is alleging OAI has a resolution of Navier-Stokes(via Córdoba and Martínez-Zoroa’s program) which maybe used info from Buckmaster and Levent Alpöge’s private Codex sessions
What the fuck is happening!!!
@Ian_Gay_briel you should write something pointing out the errors! have no idea what the general debate on this has been but the article seemed fine to me
@bwags depending on how you do the water accounting the actual number is between ~13-39,000 queries. big asterisk is agentic coding uses an insane amount of calls, but this comparison seems entirely appropriate for standard use. What context is missing?
https://t.co/ihQOXMNPCi
@ElliotGlazer the other day I thought it might be fun to draw some interesting knots and see if Gemini could find any of their invariants, complete waste of time lol, it couldn’t even correctly count their crossing number :(
@hecubian_devil@slatestarcodex@alexbronzini “Binning individual responses into income quantiles” is also just not what’s happening with the 0.98 figure. You should have someone who knows stats read through this, I suspect there are lots of minor errors.
@hecubian_devil@slatestarcodex@alexbronzini “the wealth of a group of people explains about 5% of their happiness.”
No, percentage of scale range and percentage of variance explained are unrelated concepts. As your next paragraph reports, r = .09 between log income and well being, so wealth explains 0.8% of the variance!