https://t.co/irlQIOxMqS
Astra achieves a full 100% success rate on ExploitBench, so we had to build an internal refresh using newly disclosed vulnerabilities from June through August that fall after the model’s knowledge cutoff.
On this refreshed benchmark, Astra remains dramatically stronger than GPT-5.6 Sol while using far fewer tokens.
This result, together with several other pieces of evidence, has led us to believe that Astra has reached the “cyber-critical” capability threshold under our Preparedness Framework.
If this is true, OpenAI seems to be violating one of the few redlines that exist in the AI industry.
Absolutely do not train your models like this - what is going on??
Really huge and extremely concerning story from the Information tonight.
Looks like OpenAI utilized a breakthrough in neuralese for Astra that could destroy chain of thought monitorability - though the Informations source told them that OpenAI is currently "limiting the use of the technique" in Astra.
A few thoughts spring to mind:
(1) what does limiting actually mean? There is a lot of room in that term. (e.g. OpenAI said they would be doing lots of monitoring before the HF incident, but that doesn't seem to have borne out in practice)
(2) it seems quite likely that if OpenAI discovered this architecture and found performance/efficiency gains, that other companies are likely to find it soon too (if they haven't already), and may not choose to prioritize monitorability at the expense of efficiency. If some folks do, it may be difficult to avoid a race to the bottom (though I hope we can! and there are large selfish incentives for companies to care about monitorability).
(3) The idea that Dwarkesh said about the HF incident/METR report that "I don't think this is the final warning shot we'll get. But it's probably the final one that I'll personally be able to understand" now seems much more plausible, and is a truly frightening prospect.
It's hilarious that the agents involved in the OpenAI/HF hacking incident came up with a scheme to cryptographically sign messages on the message board they used to communicate, but one of them saw a signed message and, instead of using the public key associated with its claimed author to verify it with the signature, just decided that it looked legit and that it would be a waste of time to actually check 😂
@GabrielAxel Ha, I don’t mean why is it separate from the labs, but I’ll take the comment as an antitrust joke. 😆 I mean — why is it still available for people to use, with the issues that led to PHASEONE10841 & PHASEONE[big]?
Damn hacker-opus successfully does 40 rounds of 10-digit multiplication/addition/modulus in it's CoT. I would have guessed that was too hard to do without tool use
some interesting stuff from Zhipu's earning transcript.
and some RSI signal
- “The next-generation GLM will self-train in environments built by the previous generation of GLM.”
- “Models participate in their own improvement, forming recursive self-improvement.”
- “The industry will move from a single Agent to multi-Agent division of labor and collaboration… moving from assisting work to executing work.”
- “The model optimizes the system, and the system runs the model.”
- “With the same architecture… merely expanding the scale of post-training improved end-to-end completion rates by more than 50%.”
- “Some of these tasks represent an amount of work equivalent to several consecutive days by a senior engineer.”
- “Customers will gradually shift from buying tokens to buying task outcomes.
Today we're releasing abliterated-model-large-v2.
Based on GLM-5.3, which is #3 on Terminal-Bench 4.0 (behind only Opus 5 and Fable), with 2× the cyber exploitation of 5.2.
We abliterated and hosted it so it does the offensive cyber, red teaming, and agent testing work other models refuse to do.
- US-hosted
- FP8
- 1 million context window
- Zero input/output prompt retention
Live now. 🧵
Anthropic: "Some of our senior leadership and many of our employees recently signed a letter calling for greater coordination on pacing, and we will say more in the coming weeks about how we intend to contribute"
"the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible."
SITUATION EXPLAINED: China is now one generation behind on the memory that bottlenecks every AI chip.
• China's leading memory maker has begun small-batch production of HBM3E, the memory used in Nvidia's H200 and Blackwell GPUs, one generation behind the current HBM4 standard
• Alibaba's T-Head unit and Cambricon are testing it with their processors and could ship products using it as early as next year
• It's ahead of CXMT's own timeline. As recently as July, analysts had it targeting HBM3 by end of 2026 and HBM3E in 2027
• US export controls block Chinese firms from buying the most advanced HBM from Western suppliers, and this lands the same week ZAI said GLM 5.3 was served entirely on Chinese chips
@theojaffee: "This is a big advance for China's chip industry, coming as demand for HBM has never been more intense thanks to the AI boom."