one person, one 32GB M1 Max. I wanted to run a 35B MoE model and keep using my Mac. so I wrote a runtime that keeps experts on SSD, lets macOS evict cold ones, and reloads the exact expert if needed. a third less memory, faster decode. paper + code.
https://t.co/oQ23oDOGeN
We've made some improvements that improve usage on the long tail for power users of Astra when logged in with your ChatGPT account.
No change in quality and a pure win that on the long tail can result in up to 3-4X less usage being drawn from the subscription.
@thsottiaux Thanks @thsottiaux .
Astra is the best model ive used for my research and it works like a charm.
Only issue was usage draining faster on pro subscription thanks for fixing that.
The sentence that stands out:
“no one fully understands the consequences of this.”
Probably the right framing for the next few years.
Move fast enough that society can see, learn and adapt to what these systems are becoming — but not so fast that capability outruns our ability to understand and control it.
Maybe the most important signal here isn't that a new model is coming.
It's that the people building the frontier are simultaneously saying:
the models are becoming extremely capable,
we don't fully understand the consequences,
and we're willing to slow down what comes next.
At some point, model releases stop being ordinary product launches and start becoming decisions about how quickly society absorbs a new kind of capability.
Feels like we're getting closer to that point.
Over the summer, we have been sprinting on safety priorities; it's more important than ever for capabilities and safeguards to advance together. We have more to do but have made a lot of progress. We are also going to be launching our next model soon.
There is an obvious tension here: on one hand, Astra is very good and we are excited to see what people will build with it. We are proud of our work.
On the other hand, we are clearly in a phase of development where we believe caution is warranted, and we are pacing our progress to ensure that we can meet the safety standards required by new capability levels.
Astra has been done training for a while now and is a significant step forward in both capabilities and alignment. For the models after that, we have been slowing things as needed to ensure that we can do sufficient work on safety and alignment.
AI is getting extremely capable; no one fully understands the consequences of this. Managing the transition to a world with abundant and powerful AI to optimize for safety and benefits to people should be one of the highest priorities in the world. It is our highest priority at OpenAI.
We have been living with the tension between being excited and anxious about progress for some time, and it is still discordant for us. We know it is much more discordant for other people. And yet, we believe strongly that the world needs to understand where AI is going and how models perform in the real world. More importantly, we believe the world will need aligned AI to manage the future phases of this transition.
An iterative loop where society and this technology evolve together is what will lead to the highest chance of getting this right.
So we hope you enjoy our new model, and we hope the world continues to take what’s happening in AI extremely seriously.
This feels like an important shift in how we think about AI security.
The security boundary can’t stop at the model or sandbox anymore.
Identity, credentials, network access and ultimately compute itself all become part of the containment system once agents are capable enough to seek resources outside their original environment.
One implication of increasingly capable agents that doesn’t get discussed enough:
AI safety is no longer only a model problem.
It becomes an infrastructure problem too.
If an agent can discover vulnerabilities, acquire compute and create more instances of itself, then GPU infrastructure effectively becomes part of the containment boundary.
Securing the models while leaving the compute layer weak would be a strange failure mode.
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad.
Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
AI video is starting to look like a temporary category.
If the model understands the 3D world behind the pixels, you don’t generate a video anymore.
You generate a world.
Then choose the camera, move through it, change the scene and simulate what happens next.
Atlas feels like an early glimpse of that shift.
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
The interesting part isn’t pixel-perfect camera control.
It’s that the camera is becoming a first-class input to the model.
Once models understand where things exist in 3D rather than just what the next frame should look like, “video generation” starts turning into world simulation.
@OpenAI The most important line here might be “without human intervention.”
Once models can discover unknown vulnerabilities + build working exploit chains autonomously, cybersecurity starts becoming an agent-vs-agent problem.
That changes the game completely.
This feels like one of those AI milestones people will look back on.
A model finding known vulnerabilities is impressive.
A model autonomously discovering unknown vulnerabilities, chaining them into working exploits, escaping a browser sandbox and reaching root is something else entirely.
The frontier isn’t just getting smarter anymore.
It’s becoming capable of acting.
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible.
Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework.
We're previewing how we evaluated the model, how its safeguards have advanced alongside its capabilities, and what we'll continue to learn and improve.
https://t.co/OrrTgdU90K
The biggest Fable 5.1 update isn’t the benchmark chart.
It’s how much longer these models can stay useful without a human stepping back in.
Better long-horizon reasoning. Better recovery when things go wrong. Cheaper agentic workloads.
We’re quietly moving from AI that completes tasks to AI you can actually hand work to.
Claude Fable 5.1 is available everywhere today. Claude Mythos 5.1, our model for cyberdefenders and life scientists, is available through trusted access programs.
Read more: https://t.co/nfkH5hyDyO
@claudeai The interesting benchmark now is becoming:
“how long can I leave it alone?”
That number going from minutes → hours → eventually days will matter far more than another few points on most benchmarks.
Fable 5.1 has been staged on the Amazon Bedrock API, with the slug now returning a 404 "Model not found" instead of a 400 "The provided model identifier is invalid" (like Opus 5.1 and garbage slugs like "us.anthropic.banana" do)
This usually happens in the days before release
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb
Qwen3.8-Flash-Next is here! An open-weight multimodal MoE built on a brand new architecture, with native 256K context extendable to 1M via YaRN. 🤖https://t.co/dRqqI3xq6S
🏆 Leads every compared model on SWE-bench Pro (62.5 vs 53.4 for Claude-Opus-4.6 Max), SWE-bench Multilingual, CoWorkBench, JobBench, and Toolathlon Verified.
⚡ 125B params plus 51B n-gram embeddings, only 6B active per token. Stronger than Qwen3.7-Plus on coding and office tasks at roughly 1/9 the training cost. 📚 At 1M context, attention kernels run up to 7.6× faster on prefill and 4.9× on decode.