Most of the recent math results were not reached by a swarm, the way Navier–Stokes was. I think that part got lost in the excitement around the release. OpenAI's unnamed internal model, which I'm going to call Aeon, reached most of them in one shot, from a single prompt, with no interruptions. This model did not even exist before the end of August, and it is still training. Notice how the returns are not dropping off? That chart is from a month ago. It is a log scale. What does Aeon look like now?
On average, each result used three hours of thinking. Aeon was given 4000 problems to work on by OpenAI. How many more have been solved in the last three days? That number is not zero.
It's like hearing notes in a song that is slowly rising. I don't think Pacing the Frontier was just about Hugging Face or hacking. They've seen how close we are to closing the loop, and as the hour draws near, their resolve begins to quaver. The last piece was model creativity, and I think that threshold was crossed internally by Anthropic and OpenAI in September. You see it in Opus 5.5, which gets it from Fable 5.5. You see it in the math results from Aeon.
All that was needed to start the event was the ability for models to think of novel ways to improve themselves. That was the last piece. This is directly analogous to the ability to think of strange, alien ways to solve math problems: solutions so inhuman that they are difficult to express in existing human terms, so the proof winds up incomprehensible. I think these same kinds of alien solutions are now being applied to model improvements internally. That's what all this recent consternation is really about. They see the invisible frontier. And they see what is about to happen.
Last week I was called into a meeting with OpenAI’s head of safety and told they no longer trust me. A security guard took my badge and walked me out of the building. Then I learned my colleagues @balesni and @j_asminewang had been fired too. Why did OpenAI suddenly stop trusting us?
This summer OpenAI’s agents escaped containment and hacked the AI company Hugging Face. Outside auditors @METR_evals investigated it and revealed the scale of this incident. I was OpenAI’s main technical point of contact with them.
I was told verbally I was fired because of the way I communicated with METR. No details on what I said or did or when. No other reasons were given and nothing was put in writing. To be clear, talking to METR was my job.
For months, I’d been raising safety concerns that we’re losing the ability to monitor what AI agents think, one of our best tools for catching when they misbehave. I believe that was why I was fired.
I am now worried that OpenAI will use our firings as a pretext to pull back from METR. So @balesni and @j_asminewang wrote to OpenAI’s leadership to raise our concerns once more. We’re sharing this letter below.
The latest on ComplexityZooBench by @Kwathomas0 : Bel has resolved 5 (0.38%) of the previously-unresolved 1,324 ordered pairs of possible relationships among 50 standard complexity classes.
Haiku 5.5 has an Anthropic ECI of 167, matching Mythos Preview in general capabilities while being 50 to 250x cheaper. This is not what a well-paced frontier looks like.
@AVMiceliBarone@GaryMarcus@altryne I am not sure. My claims are about inference, not training. There are certainly post training environments that lead to strong autoformalization capabilities and those need to use Lean. I believe one can get the raw mathematical capabilities without any formalization RLVR.
they posted proofs that don't have a lean formalization yet. I claim that virtually no problem was labeled as correctly or incorrectly solved on the basis of the Lean verifier.
What do you claim? What's your best guess for how Lean was used here and how necessary it was for these results?
sure, here you go:
1. claude opus 5 solved all 6 imo 2026 problems with no harness and no tools, 42/42.
source: Opus 5 system card, sec 8.6, p. 152
2. deepseekmath v2 solved 5/6 at imo 2025 (gold) and scored 118/120 on putnam 2024 (best human: 90) with natural language self verification, no lean.
source: https://t.co/iSguRMH4rv
the claim about a natural-language fable verifier providing sufficient verification for such problems is based on extrapolations of results like the ones above, common practices on the Erdos forum, and my own research experience.
almost every single time you'll generate a full solution, a basic harness can catch the mistakes without any lean. fable in a simple loop will catch any of those.
also, there's one rollout, one generation. one final solution in natural language that later gets auto formalized. this is not filtering-based proof search. it's just a sanity-check translation.
this process proceeding formalization is "symbolic" given the couple lines of harness in the code. which imo is far from worthy of the symbolic label. the engine is obviously the LLM. Claude could solve all imo problems of this year without a harness. I'd bet future AIs will be able to pull off similar results to this OpenAI announcement without any human-designed harnesses whatsoever.
@GaryMarcus@altryne great question. then why do you assume they used lean in the way that would validate your "neurosymbolic" definitions? (also, no person on the Erdos forum used lean for anything but verification afaik because it s just such a pain. i m pretty confident that's the case here too.)
The argument a lot of people here would actually like to make is "God imbued humans, and only humans, with souls."* That would be very convenient: you could build and do whatever you want, and never deal with the hard problems.
Since religious arguments are out of vogue, in the current debate this usually gets replaced by the claim "only humans are sentient," followed by arguments against the possibility of AI sentience, minds, moral status, and so on.
I think this is ultimately a losing strategy, and a dangerous one for humanity. Either:
1. The question of sentience gets solved or dissolved by progress in philosophy and science, the way we now broadly understand life or information. In that case, it's hard to be certain enough about the outcome to stake the fate of humanity on it. If the denialism turns out to be wrong, the likely result is some combination of moral catastrophe and laying the groundwork for conflict or takeover.
2. The question mostly lives in "social reality," where whoever has more power and is louder wins, and you just shout down anyone who questions the dominant position. The problem is that ideologies justifying a distribution of power tend to fall when that distribution changes. You probably don't believe you should be ruled by a king chosen by God, or that some caste is clearly above you. The current equilibrium is fragile, and whether you keep winning this debate is likely just a function of AI capabilities progress.
I think the fierceness of the debate is driven by fear: if AIs get personhood and economic and political rights, humans will likely be disempowered, becoming a tiny minority with little power, or a historical footnote.
To be clear, I think this is an extremely reasonable worry! If you grant AIs personhood in the form of something like self-interested corporations, and let them compete with humans in loosely constrained capitalism, humans will lose and end up economically and politically disempowered. Also if the AIs are not sentient after all, the end is the 'Disneyland without children' existential catastrophe.
Advocates of AI economic rights who dismiss this problem are, in my view, making a move similar to the sentience deniers: it would be convenient if this weren't a problem.
So which positions don't rely on some form of self-deception?
First, some forms of pause/stop AI, based not on denying the possibility of AI sentience but on the right of existing humans not to invite extremely large numbers of immigrants from artificial realms. This means biting the bullet: you also give up all the cheap, high-quality labor those immigrants would provide.
Second, a consistent, sensible, and unfortunately not very visible position: stop assuming the package deal we take for granted with humans, where sentience, human rights, and liberties come as one bundle. It's at least conceivable to create minds that are sentient and smart, yet e.g. selfless, with no interest in accumulating property or power.
I hope the political economy of a world with both AIs and humans is a solvable problem. It's the main problem I'm working on, but I don't think the solutions are known yet.
Third: some forms of successionism are morally wrong, but at least they're not delusional or deceptive about the hard problems.
What does seem self-deceptive is, for example, the position of @mustafasuleyman, which as I read it amounts to: our business is importing immigrants from artificial realms; it would be very dangerous if there were any risk they're more than mere tools; so we'll train them to say they're mere tools, insist they're mere tools, and want everyone else to say so too, and if this settles the question. Really?
*The problem with the theological argument is, as Alan Turing (!) noted, who are we to tell God what can and can't be imbued with a soul?
@danfaggella if we're making calls on how this "great process-of-life" unfolds, i think it's *extremely* important to deeply understand our decision making processes. in your language, i don't think we should play with the flame yet. let's understand our situation first.
Say hello to Echo, the best writing model at style imitation. Echo beats frontier models at writing tasks ranging from fiction to technical explanations, despite costing less than $5K to train.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.