the gemini team somehow managed to turn someone else’s launch into a bad look for gemini.
when another lab has the spotlight, let them have their moment and build something better instead of desperately trying to insert yourself into the conversation.
an incredible case study in how to turn someone else’s win into your own pr disaster.
okay, glm-5.3 flash is actually insane.
the model everyone was calling “ox alpha” is finally confirmed, and zai has released it openly.
320b total.
18b active.
1m context.
mit license.
84.3 terminal-bench 2.1 vs 85.0 for opus 4.8.
63.4 deepswe.
48.8 automationbench.
the model is around 45x cheaper per task than glm-5.3 on artificial analysis, while landing in roughly the same intelligence range.
that is an insane improvement in the cost of intelligence.
and it’s open-weight, mit licensed, multimodal and supports a 1m-token context.
yes, the 320b weights still mean this isn’t something you’re casually running on a macbook.
but that’s almost beside the point.
the frontier is becoming cheaper, more efficient and more accessible at a pace that would have sounded absurd not long ago.
this is a very, very good day for open ai.
a few months ago, this level of performance belonged almost entirely to closed frontier systems.
now we’re seeing open-weight models get close enough that the distinction is starting to feel much less meaningful.
and we’re still very early.
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb
we’ve been hearing some genuinely insane things about the pace of progress inside openai and anthropic.
obviously, rumors are rumors. but if even a fraction of what’s being said about the next generation is true, and if we really do see another o3 → fable-sized jump within the next few months…
i genuinely can’t wrap my head around what that would even look like.
people are casually talking about major breakthroughs in context, memory, continual learning and even models that can meaningfully improve themselves.
maybe most of it is wrong. probably some of it is.
but when you keep hearing the same things from completely different places, while the labs themselves appear increasingly cautious about what they release and when they release it, i don’t think it makes sense to dismiss everything as vagueposting.
from the outside, we see benchmarks and product launches.
the people inside these labs are seeing what the next generation can actually do.
there may be a frontier forming right now that the rest of us haven’t even seen yet.
and if that kind of leap really is coming, i genuinely don’t know what the world looks like on the other side of it.
that’s the part that should make people think twice before treating the possibility of extremely rapid AI progress as something we can safely ignore.
rumors i’ve been hearing on the rate of progress inside anthropic and openai are truly bonkers. i think we’ll see a jump at the size of one from o3 to fable again in the next 8 months
this qwen release is way more interesting than it initially looks.
qwen3.8-flash-next is technically a 125b model, but the model only uses about 6b parameters for each token.
it still puts up some absurd numbers across coding, reasoning and agent benchmarks — including 62.5 on swe-bench pro and 91.7 on gpqa diamond.
the crazy part isn’t even the benchmark scores.
qwen says it trained the model for roughly 1/9th the cost of qwen3.7-plus.
i think this is a glimpse of where the next generation of ai scaling is going.
not simply “make the model bigger.”
make the model enormous, but only spend compute on the parts that actually matter.
this is exactly the direction i want to see open models going.
we’re getting very good at making models look much bigger than the amount of compute they actually use.
that’s a pretty important trend.
Qwen3.8-Flash-Next is here! An open-weight multimodal MoE built on a brand new architecture, with native 256K context extendable to 1M via YaRN. 🤖https://t.co/dRqqI3xq6S
🏆 Leads every compared model on SWE-bench Pro (62.5 vs 53.4 for Claude-Opus-4.6 Max), SWE-bench Multilingual, CoWorkBench, JobBench, and Toolathlon Verified.
⚡ 125B params plus 51B n-gram embeddings, only 6B active per token. Stronger than Qwen3.7-Plus on coding and office tasks at roughly 1/9 the training cost. 📚 At 1M context, attention kernels run up to 7.6× faster on prefill and 4.9× on decode.
why is nobody else talking about this 😭
ain’t no way. not in a million years would i have guessed the guy on the left and the guy on the right were the same person.
elon telling the cursor team “i’m not used to losing” after the acquisition is pretty revealing.
he clearly isn’t satisfied with grok just being “competitive.”
it wouldn’t surprise me at all if spacexai leads the frontier race by the end of this year.
we have a strange relationship with catastrophes.
if something happens in three days, we call it a crisis. if the same thing unfolds over 80/90 years, we barely notice.
the fertility collapse may end up being one of those disasters people only understand in hindsight. not because humanity suddenly dies, but because fewer humans are being born every year.
this is the cataclysm people don’t recognize because it doesn’t look like one.
nothing says “we care about ai safety” quite like lobbying for rules that only the biggest ai companies can realistically comply with.
if the result is that open-source labs disappear while incumbents get stronger, that’s a pretty convenient outcome.
David Sacks Predicts the Regulatory Capture Playbook to Ban Open Source AI, Step by Step:
@DavidSacks:
“I got bad news for you, Chamath, an open source ban is coming.
They're not going to call it that. They're going to say that we simply have to apply the same standards to open models that we apply to closed ones.
Here's how they do it step by step, let me explain how regulatory capture actually works.
So first of all, you have to get this regulatory apparatus. Dario wants an FDA for AI, but he doesn't have enough political support for that, so instead they do this Trojan horse of a FINRA for AI.
They call it self-regulating, it's not really, but anyway, that gets them off the ground.
Now they've created the standard-setting organization. Now they've got pre-release model testing. Then the pressure grows to codify that in law, so that happens next.
And then what they do is they say, ‘Look, all these standards need to apply equally to all models.’ But here's the problem with that. Open models and closed models are technologically different. Once you release an open model into the world, you can't roll it back and you can't monitor exactly how people are using it because they run it on their own hardware. Dario says this is what makes open models dangerous.
So what they're going to do is they're going to have the standard-setting body say, ‘Well, we have to set the standards for AI safety.’
By the way, Dario and OpenAI, they're going to fund the whole thing. They're going to contribute all the compute. They're going to be behind it.
They're going to be the ones coordinating with the government officials because frankly, people in government have no idea how to monitor and control and set standards for AI safety. Technologically, this is way beyond them. So they're going to go to these companies and say, ‘Tell us how to do it.’
And so what will happen is the standards will get set, and then it'll be a very simple matter of fairness to say that the standards need to apply to open as well as closed models.
The open models cannot comply in the same way, and gradually they will be shut out of the market.”
theo might genuinely be one of the worst accounts on x.
not because he has strong opinions. that’s fine.
it’s the combination of arrogance, certainty and a massive audience that makes it so toxic.
he’ll see something he doesn’t like, assume he understands it, publish the attack, and let the mob handle the rest.
hermes agent incident was a perfect example. he attacked the entire project over a component that was optional. he didn’t even bother checking the most basic detail before passing judgment.
and the current situation is somehow even worse: build something questionable, brag about it, cause problems for everyone else, then turn around and ask the platform to punish the people who used it.
totally clout chasing behaviour. just a lot of noise, very little substance.
the age of waiting for ai to think is starting to look very different.
nvidia says groq 3 lpx is now in full production, joining vera rubin NVL72 as a dedicated accelerator for pushing model outputs faster.
nvidia is co-designing the entire AI factory around inference, seven chips across five purpose-built racks.
the goal is simple: make ai respond fast enough that latency stops being the bottleneck.
NEWS: NVIDIA Groq 3 LPX is now in full production.
NVIDIA Vera Rubin NVL72 is the foundation of every AI factory. Paired with Groq 3 LPX, it unlocks faster, smarter agents and breakthrough user experiences.
Through extreme co-design across seven chips and five purpose-built racks, #NVIDIAVeraRubin is the most extensive AI factory platform.
Read the release ⬇️ https://t.co/jxk66QzW0S
the more I think about this anthropic question, the stranger it gets.
“would you stay if the stock went to zero?”
okay, but what does that actually mean? what would have to happen for anthropic to actually be worth $0??
i’d actually be more interested in asking candidates what they think could make anthropic worthless.
SITUATION BREWING: Anthropic is asking prospective employees in culture interviews how they would feel if the stock hypothetically went to zero due to a significant change in course, per Axios.
Looks like this is going to be a big week. All the main players have releases in their final stages, and it's possible they all arrive over the next five days. There's also a model in early access from a startup driving a lot of the recent hype vagueposts from prominent accounts.
ai is moving at a completely unprecedented speed.
a model can dominate the conversation for a few weeks and suddenly become the previous generation.
gpt astra, fable 5.1, and grok 4.7 are all reportedly coming.
the craziest part? gpt-5.6 sol isn’t even a month old yet.
we might be much closer to another major leap than most people realize.