I make security interesting. Award-winning HRM awareness & culture change leader. Also, I break AI so you don’t have to. DM for existential dread & consulting.
🤔 If you showed someone from six years ago Astra, they might agree. But we’ve learned more about what AGI really means since then (it hasn’t just been needlessly moving goalposts). Maybe ace HLE and achieve an @ArtificialAnlys Intelligence Index of 90+ to end debate? Soon. 🧠
@ChaseLochmiller@OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations @OpenAI team.
400K GPUs coming online next.
Today we're releasing data on models accelerating research at OpenAI.
Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same.
https://t.co/iLKbrLcBAI
GPT-6 Astra is AGI.
91.8% vs an 80% human baseline on SpatialBench. It’s not just matching humans, it’s beating us at spatial reasoning, one of the last areas where humans still had a clear edge.
This feels like a line being crossed.
Here’s what I’ve learned from inside OpenAI.
The next major model after Astra (not updates like Astra 6.1) will be released as AGI. It will be able to perform actions in real time, including playing real time games. Its capabilities will extend beyond next-word prediction, with humanlike general understanding, expert-level performance across all fields, and some capacity for recursive self-improvement.
The release is expected toward the end of 2026, potentially in early or mid-November.
Internally, OpenAI has already reached the early stages of recursive self-improvement, and the impact is already substantial. 5.6 Sol helped train Astra, and Astra helped train Doug and Bel. Doug, the larger and more capable pretrained model, is now helping train its successor.
That successor is expected to take recursive self-improvement further, with the ability to outperform all humans at every task, including those requiring real-time action. The timeline I’m hearing for that is mid-2027, with late 2027 also possible.
My first "holy shit" moment with GPT-6 Astra:
I asked it to create a world in Unreal Engine, and fill it with humans (each an Astra-powered agent) who all have to work together to survive.
A day later, I was in my bedroom and heard voices coming from the living room... I thought someone was in my apartment.
I walked out, honestly a little scared.
It was the Astra agents. They'd started talking to each other.
Fucking crazy.
Here's a brief clip (obviously not 100% perfect yet, but still, insane. sound on!):
GPT-6 Astra is here.
We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building.
We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more.
It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait.
It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.
This is exactly why we open source!
You build the model, put it in people’s hands, and watch them take it somewhere magically unexpected.
MiniMax H3-Max just getting started!
One of the many things I love about GLM is that it has the lowest hallucination rates.
GLM 5.3 Flash has by far the lowest hallucination rate among the frontier models. GPT 5.6 and Opus 5 hallucinate ~3x more often.
We have a limited window to strengthen cyber defenses, and together with organizations including @AnthropicAI, @awscloud, @Google, @Microsoft, and @Oracle, we're calling for a global effort to give defenders the tools, resources, and support to protect the infrastructure we all depend on.
If we act decisively, we can turn today's AI advances into lasting improvements in security and make our digital world safer for everyone.
https://t.co/f33JRVCiJb
this is a critically important moment for cyber defense with AI; there is not much time to act.
we are happy if you want to work with us or any of our competitors or partners, but please take this moment seriously.
only an urgent and intense collective response will work.
@pollenrobotics Make a version big enough that it can climb up and down stairs, and give it a tall, multi-joint grabber arm capable of lifting 5 pounds (pick up laundry, put it in a hamper/washer/dryer) and towing/dragging 15 pounds (drag laundry bag up or down stairs) and I would pay 3x as much
everyone assumed it would be google that would unseat or threaten openai’s lead on models but the one that gave them a run for their money was an (at the time) smaller company called anthropic
it is now in the interests of both incumbents to say, and possibly true: RSI is at hand, join or become irrelevant. but i would wager while the fortunes and misfortunes of research could shift the immediate race in either’s favor or a third large company like meta or xai, history is always longer than people account for
first, as A\ proved, huge amounts of compute are more elastic than some previously thought. there is some inference bid for which major giants will give up some of their frontier training ambitions and rent you parts of their compute buildout (google, spacexai etc )
a company betting on something radically different in research or execution could gain a major foothold especially as so much of the stack becomes autocatalytic— in other words you could use a frontier model to help build the parts of your platform you’re not good at. if you’re not constrained by American law you can even distill Claude and GPT to close the gap
small experimental labs are likely a bigger threat to OpenAI and anthropic than large companies
it is game theoretically hard to not serve your frontier model - if your competitor is willing to launch then you may lose a huge chunk of your revenue. if safety concerns supersede that, all of this potentially flips if everyone at the frontier uniformly keeps the strongest models hidden or nerfs their AI R&D capabilities for many months behind the frontier. that may lead to closed loop RSI & runaway advantages. even a huge breakthrough in compute efficiency won’t matter if the labs have private engines of autocatalysis
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency.
What's new: 🥳
- Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4.
- Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks.
- Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI).
- 262K native context, extensible to 1M with YaRN.
We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀
We can't wait to see what you build with Qwen3.8-Flash!👀👇
- Blog: https://t.co/M5hYypFLgJ
- Technical Report: https://t.co/IF0gObIkQO
- Hugging Face: https://t.co/6ow8QVAABt
- ModelScope: https://t.co/tDOn2jNuFG
Nvidia has lost it's moat 🤯
Microsoft has open-sourced a 1-bit inference framework that runs massive 100B parameter models directly on your CPU without GPUs.
82% less energy. 6x faster inference. 100% open-source.
ultrafast inference reveals new threats. worth thinking about how quickly misaligned frontier class models running 50x faster could infiltrate systems and so on— far too quickly for human responders to stay abreast. you need autonomous detection and shutdown, not just monitoring