There is a heated debate right now. Tristan Buckmaster says OpenAI told him its AI had solved the Navier–Stokes Millennium Prize Problem.
That claim remains unverified, and the discussions around it have erupted into a dispute over scientific credit.
Here is what happened and what we know so far:
Several different things are getting mixed together. Here is what the available evidence actually says.
Building on work by Córdoba and Martínez-Zoroa, Alpöge and Buckmaster used Claude and OpenAI models to establish how initially smooth fluid flows can develop singularities under smooth external forcing. They released papers and Lean formalizations, and Terence Tao praised the advance.
Their results cover Euler and related equations. They do not solve the Navier–Stokes Millennium Problem, which additionally involves viscosity. A suitable extension could resolve it, even with smooth external forcing.
Then comes the disputed part.
According to Buckmaster, OpenAI told him on September 6 that an internal model had already produced a proof for forced Navier–Stokes, potentially taking that final step.
He says OpenAI’s effort began after information about their progress reached the company. He also disputes the initial description of minimal human involvement, reporting that the discussions revealed a team effort, multiple attempts, and substantial compute.
Buckmaster further says publication proposals included having him write up OpenAI’s result without Alpöge as a coauthor, with Alpöge’s employment at Anthropic cited as an obstacle.
He reports being asked, “Why would you ruin your career?” after saying he would make the circumstances public.
Bubeck has publicly rejected allegations against him as “false and inflammatory.” These accounts are contested. The social posts do not establish what happened in the private discussions.
There is also no established evidence that OpenAI trained on their private Codex sessions or stole their drafts. Buckmaster raises the question but explicitly writes: “I do not know whether our data was used.”
tl;dr
Alpöge and Buckmaster used AI to prove new results about how smooth fluid flows can develop singularities, advancing a research program relevant to the Navier–Stokes Millennium Problem. Their published work does not solve that problem, but Buckmaster says OpenAI privately told him its internal model had taken the remaining step. He has not seen that proof, and the claim remains unverified.
But rumors are that OpenAI actually fully solved that milleniun problem.
I deleted my previous post. I'm just going to post the entire statement page by page. It's a complicated situation.
Statement: https://t.co/QuiFawTO5n
https://t.co/2WnmsM5Vb5
https://t.co/LwPwCdQiJo
https://t.co/J5bDTrWDnE
https://t.co/MKYMXrrZPy
DripStocks for the Base Builder Quest @buildonbase
Stream mock tokenized-stock payroll on Base Sepolia. Create a secure claim link, bind it to a recipient wallet, claim, then withdraw vested AAPLc anytime.
Live: https://t.co/ppzoxDblQw
Code: https://t.co/oBk87vdvMd
Never gonna give you up
Never gonna let you down
Never gonna run around and desert you
Never gonna make you cry
Never gonna say goodbye
Never gonna tell a lie and hurt you
Thanks for reading. We will do a global reset of the usage for all paid subscriptions so that you can keep enjoying Astra after burning through all of it doing fun 3D modeling in blender. The work week is about to start.
Lands around 6pm PST today.
ChatGPT Work can now pick up on what makes your writing sound like… you. Your favorite phrases. Your very specific sign-off. your capitalizations quirks.
Connect the tools you use every day, like Gmail, Google Drive, Slack, and SharePoint, and ChatGPT Work will learn your writing style from your emails, messages, and files—and carry it into whatever you write next.
Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5
Changelog (Index v4.2 → Index v4.3):
➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench
➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark
We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation
Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%
Detailed changes:
➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon
➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6
Key results:
➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47)
➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36)
➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
To calibrate you all on which reasoning effort to use for Astra, know that GPT-6 Astra on low performs better than GPT-5.6 Sol on high.
If you were using high reasoning efforts with Sol and were happy, I suggest you move down to low or medium for Astra.
@ChaseLochmiller@OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations @OpenAI team.
400K GPUs coming online next.
OpenAI have reached their goal of an automated research intern, and now say they are on track for a fully autonomous AI researcher 18 months from now: in March 2028.
Key Events This Week:
1. US Markets Closed, Labor Day - Monday
2. US 10Y Note Auction - Wednesday
3. August PPI Inflation data - Thursday
4. August Existing Home Sales data - Thursday
5. August CPI Inflation data - Friday
6. September MI Inflation Expectations data - Friday
7. September MI Consumer Sentiment data - Friday
This marks the final week of inflation data before the September Fed meeting.
@DanDr1s Qwen 3.8 Max in the same tier as Opus 5, Sonnet 5 in the same tier as Muse Spark 1.3, and Kimi K3 being C tier.
This list is genuinely beyond laughable