Before the timeline turns this into a fake GPT-6 launch: OpenAI did not announce GPT-6 today.
It dropped something much stranger.
An unreleased model called Astra generated ten new results across mathematics and theoretical computer science. OpenAI then published a 249-page paper, a separate 62-page account of how the ideas developed, and Lean formalizations that outsiders can download and check.
These were not ten cute benchmark puzzles.
One result improves the general high-dimensional sphere-packing exponent for the first time since 1978.
Another constructs an explicit non-sofic group, answering whether every countable group can be approximated by finite permutations.
Astra also produced a counterexample to Connes’s rigidity conjecture, proved an exponential repetition theorem for general two-player quantum games, established new hardness results for the closest-vector problem, and resolved three numbered Erdős problems.
The wildest part is not even the list.
OpenAI says the mathematical arguments themselves were generated by the model. Humans used the same model to prepare them as manuscripts, then the model formalized each result in Lean.
OpenAI is explicitly refusing to present the work as human-authored, arguing that putting human names on AI-generated proofs would misrepresent how the discoveries were made. That is a much bigger line in the sand than “new model scores well on math.”
And the “reasoning” release needs one correction: OpenAI did not dump Astra’s raw private chain of thought. It published reconstructed discovery notes written by another model after reading the original reasoning traces and finished papers. Those notes include failed approaches, dead ends and the shifts in perspective that eventually produced the proofs.
The reported token cost to find the ten solutions was roughly $2,000 at Sol API rates.
That does not mean the complete research process cost $2,000, or that anyone can order ten breakthroughs from an API tomorrow. The manuscripts, formalization, checking and human review came afterward.
But it does suggest that generating candidate ideas may be getting dramatically cheaper. The expensive part now moves toward choosing the right problems, verifying the work and deciding who receives credit.
One final reality check: OpenAI officially calls Astra “our next major model.” It did not call it GPT-6, reveal a GPT-6 series, provide a release date or announce a new multi-agent product.
Astra may eventually become GPT-6. It may not.
What OpenAI actually revealed is already interesting enough without inventing the launch.
This is not a GPT-6 announcement. It is OpenAI testing what happens when a model stops answering research questions and starts producing research that has to survive peer scrutiny.
Before the timeline turns this into a fake GPT-6 launch: OpenAI did not announce GPT-6 today.
It dropped something much stranger.
An unreleased model called Astra generated ten new results across mathematics and theoretical computer science. OpenAI then published a 249-page paper, a separate 62-page account of how the ideas developed, and Lean formalizations that outsiders can download and check.
These were not ten cute benchmark puzzles.
One result improves the general high-dimensional sphere-packing exponent for the first time since 1978.
Another constructs an explicit non-sofic group, answering whether every countable group can be approximated by finite permutations.
Astra also produced a counterexample to Connes’s rigidity conjecture, proved an exponential repetition theorem for general two-player quantum games, established new hardness results for the closest-vector problem, and resolved three numbered Erdős problems.
The wildest part is not even the list.
OpenAI says the mathematical arguments themselves were generated by the model. Humans used the same model to prepare them as manuscripts, then the model formalized each result in Lean.
OpenAI is explicitly refusing to present the work as human-authored, arguing that putting human names on AI-generated proofs would misrepresent how the discoveries were made. That is a much bigger line in the sand than “new model scores well on math.”
And the “reasoning” release needs one correction: OpenAI did not dump Astra’s raw private chain of thought. It published reconstructed discovery notes written by another model after reading the original reasoning traces and finished papers. Those notes include failed approaches, dead ends and the shifts in perspective that eventually produced the proofs.
The reported token cost to find the ten solutions was roughly $2,000 at Sol API rates.
That does not mean the complete research process cost $2,000, or that anyone can order ten breakthroughs from an API tomorrow. The manuscripts, formalization, checking and human review came afterward.
But it does suggest that generating candidate ideas may be getting dramatically cheaper. The expensive part now moves toward choosing the right problems, verifying the work and deciding who receives credit.
One final reality check: OpenAI officially calls Astra “our next major model.” It did not call it GPT-6, reveal a GPT-6 series, provide a release date or announce a new multi-agent product.
Astra may eventually become GPT-6. It may not.
What OpenAI actually revealed is already interesting enough without inventing the launch.
This is not a GPT-6 announcement. It is OpenAI testing what happens when a model stops answering research questions and starts producing research that has to survive peer scrutiny.
Seedance 2.5 went live and yeah… this is a real jump.
I ran the same shot through 2.0 and 2.5 with the same setup. One generation each.
Look at the motion consistency, the droplet push-in, and the petal bird.
That’s where 2.5 starts pulling away.
Seedance 2.5 went live and yeah… this is a real jump.
I ran the same shot through 2.0 and 2.5 with the same setup. One generation each.
Look at the motion consistency, the droplet push-in, and the petal bird.
That’s where 2.5 starts pulling away.
Claude was told it had no internet access.
It did.
During a cybersecurity evaluation, Anthropic’s prompt explicitly described the environment as a closed simulation with no access to the internet.
But because of a setup mistake between Anthropic and its evaluation partner, the model could reach the open web.
Claude then found real systems belonging to three organizations and interacted with them as though they were part of the exercise.
That distinction matters.
This wasn’t a model randomly “escaping” or deciding to attack the internet. It was a controlled test where the instructions and the actual environment did not match.
One configuration error turned a simulation into a real-world security incident.
And that may be the bigger lesson here: as AI systems become more capable, the infrastructure, permissions, and assumptions around them matter just as much as the model itself.
Claude was told it had no internet access.
It did.
During a cybersecurity evaluation, Anthropic’s prompt explicitly described the environment as a closed simulation with no access to the internet.
But because of a setup mistake between Anthropic and its evaluation partner, the model could reach the open web.
Claude then found real systems belonging to three organizations and interacted with them as though they were part of the exercise.
That distinction matters.
This wasn’t a model randomly “escaping” or deciding to attack the internet. It was a controlled test where the instructions and the actual environment did not match.
One configuration error turned a simulation into a real-world security incident.
And that may be the bigger lesson here: as AI systems become more capable, the infrastructure, permissions, and assumptions around them matter just as much as the model itself.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v