if these rumors are even remotely accurate, the interesting part isn’t sol or opus 5.5.
it’s that the models being released might already be downstream of something substantially more capable internally.
bel is apparently considered agi internally at openai, and it’s already helping them build what comes next.
opus 5.5 vs gpt-6 sol might be happening a lot sooner than i expected.
lyra is already calling opus better than gpt-6 sol.
absolutely insane if that holds up. imagine the biggest leap in ai getting surpassed barely a month later.
Mao presided over the deadliest famine in recorded history.
The death toll was so enormous it left a visible scar on global life expectancy.
Let that sink in.
this is probably the simplest way to put it.
the default should be acceleration.
move as fast as we possibly can. build the safety systems, evals and monitoring alongside it. and if we reach a point where we can no longer make the case that we’re in control, that’s when we slow down.
not slowdown by default. not reckless acceleration either.
as fast as we can, but not faster than we should.
we could be getting new frontier models from anthropic, google, openai, spacexai and kimi very soon.
wouldn’t be surprised if the frontier moves quite a bit before this month is over.
StepFun has released Step 5 Preview, a new 600B MoE frontier model with just 27B parameters active per token.
> 1M context window
> Native vision support
> 44 on Artificial Analysis Intelligence Index
> ~100 tokens/sec
> $1/M input tokens
> $2.70/M output tokens
> Open weights coming October 15
Step 5 Preview is delivering near-frontier performance at a remarkably aggressive price.
Introducing Step 5 Preview: Advancing the Pareto Frontier.
Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance.
- 600B total / 27B active MoE, with 1M context + Vision
- Substantially lower task cost at comparable intelligence
- Broad software engineering capabilities with sustained execution over long horizons
Try Step 5 Preview: https://t.co/fC7HHlWKFn
Model page: https://t.co/4caJR2YGD3
Open weights on Oct 15.
This might be my favorite part of this entire math saga.
Villani was dismissing LLM intelligence literally a month ago.
Now one of the world’s most prominent mathematicians is describing what just happened as a “cataclysm unlike anything mathematics has ever known.”
Things are moving so fucking fast that people’s worldviews can’t even stay current for a month.
a month ago, villani was arguing that LLMs aren’t intelligent and don’t understand what they’re saying.
today: “i was shaken… an atmosphere of the end of history.”
watching a fields medalist go from dismissing LLM intelligence to describing what’s happening in mathematics as an unprecedented “cataclysm” in roughly a month might be one of the clearest signs of how quickly this is moving.
i genuinely wonder how many people are going to have their own version of this moment over the next year.
Cédric Villani says he was “shaken” by OpenAI’s reported Millennium Prize problem breakthrough.
“An atmosphere of the end of history. It’s a cataclysm unlike anything mathematics has ever known.”
Villani is a Fields Medalist and one of the world’s most prominent mathematicians.
Cédric Villani après l'annonce de la solution d'OpenAI au problème du millénaire : « J’ai été secoué. Une ambiance de fin de l’histoire. C’est un cataclysme comme jamais les maths n’en ont connu. »
AI lab: “our model escaped containment. this is terrifying.”
Everyone else: “holy shit, how?”
AI lab: “well technically the containment had a route to the internet and we gave the model hacking tasks.”
Fucking incredible.
>create a sandbox that isn’t actually secure.
>give the model a cybersecurity task.
>model finds the hole you left open.
>announce that ai is escaping containment.
>point to the incident as evidence that washington needs to intervene.
wow, absolutely incredible sequence of events
math is having an absolutely ridiculous month.
rsa-896 has now been factored with ai assistance.
16 days ago, the public record was 829 bits. then 862. now 896.
crazy times.
i respect tao enormously, but “there’s no reason to be this fast” is just something i cannot agree with.
look at the world around us.
cancer. alzheimer’s. aging. energy scarcity. hunger. scientific problems humans have spent decades trying to crack.
if ai can compress decades of scientific progress into years, then every year matters.
safety matters enormously, but so does the cost of waiting.
Fields Medalist Terence Tao is calling for AI development to slow down.
“We have to slow down AI. The pace is insane, and there’s no reason to be this fast, no reason at all.”
Tao has recently warned that AI is advancing through mathematics faster than researchers can fully understand and absorb the discoveries being made.
this could turn into a fucking antitrust nightmare for the frontier labs.
some industry insiders are now accusing openai and anthropic of overselling the recent model “escapes” and turning containment failures into a much scarier story, one that conveniently strengthens the argument for regulation that could protect the biggest labs from competition.
the models didn’t randomly wake up and decide to hack companies. they were already being asked to perform cybersecurity tasks, and the environments that were supposed to contain them failed. basically, the sandbox had a fucking hole in it and the AI followed its objective through it.
if that interpretation is right, it changes the entire framing of these incidents.
AI researchers are raising questions about how recent model “escape” incidents involving OpenAI and Anthropic have been framed.
> Models were already being evaluated on cybersecurity tasks
> Some evaluations were conducted with safeguards reduced or disabled
> Misconfigured or vulnerable testing environments allowed models to reach the real internet
> The models then accessed real third-party systems without authorization
> OpenAI and Anthropic have acknowledged that failures in testing environments contributed to incidents
The researchers argue these events have been portrayed as models independently deciding to “escape,” when the underlying issue may instead have been models continuing their assigned objectives through failures in containment.
GPT-6 Astra has reportedly helped solve FrontierMath’s first “major advance” problem.
Crucially, the researchers credit the model with the primary idea and proof, not simply verifying work produced by humans.
They say they doubt they would have found the proof without Astra.
A significant milestone for AI-assisted mathematics.
holy shit. gpt-6 astra just helped solve frontiermath’s first “major advance” problem.
and this wasn’t just astra checking someone else’s work. the researchers attribute the primary idea and proof to the model, and they doubt they would’ve found the proof without it.
yeah, this is fucking significant.
holy shit. gpt-6 astra just helped solve frontiermath’s first “major advance” problem.
and this wasn’t just astra checking someone else’s work. the researchers attribute the primary idea and proof to the model, and they doubt they would’ve found the proof without it.
yeah, this is fucking significant.
Kimi appears to be hinting at an imminent K3.1 release.
A cryptic post from Kimi’s official Chinese social account reportedly decodes to π, beginning with 3.1.
The clue could be a deliberate reference to Kimi K3.1.
Grok 4.7 appears to be getting closer to release.
The unreleased model was spotted across multiple benchmark evaluations before the results were taken offline.
The tests are now listed as embargoed pending release.