We went quiet for a few weeks. Reason: We rebuilt everything. Our own stack, same mission — open-source AI measured, not hyped. V2 is in final gate checks right now: every ranking is math you can verify, every number traces to a source.
What should a directory have to prove before it ranks anything?
WOW!
GROK 4.5 PRICES ARE NEAR OPEN SOURCE HOSTED PRICES!
This will move many that have token cost fatigue in corporations.
But that is only half the story: it is ~4.2 more efficient using far less times per turn.
When you love building it's so easy to overbuild which pulls away from what you really set out to build in the first place. Convolutes the product beyond recognition of the primary objective.
I always feel like I should deliver more value, more value, more value but it becomes a build swamp with little value. lol
Ok so I rebuilt OSS AI Hub from scratch.
The old one had a stack builder, discussion boards, all this stuff I was sure people wanted. Turns out, no. me included.
So, I cut it and kept the one thing that mattered: a read on which open-source AI tools are actually moving. 2,477 of them, ranked by real GitHub velocity, refreshed nightly. no featured slots, no pay-to-rank — just the numbers.
Plus, a plain-English glossary and two calculators that actually answer something: a VRAM one (the right GPU to buy or rent for your model) and a cost one (what your traffic runs across the major LLM APIs).
live now: https://t.co/tM3f9ZOaDR
what's something you were proud of, then killed?
"I HELP [AUDIENCE] ACHIEVE [RESULT]" is a costume. The account that wins is the one that took it off. You spent 15h learning to sound like everyone who took that course. The thing nobody can copy is you being you, out loud, figuring it out. That's the whole moat and you do it perfectly.
A new open model drops every few days, each one “closing the gap” on coding or agents.
Easy to feel behind if you’re not testing the latest thing the day it lands.
Here’s what I keep noticing: it matters less than the timeline says it does.
The people quietly shipping in production aren’t chasing every release. They picked a model that fits their hardware and constraints, then poured the work into everything around it — evals, recovery, scoping. The unsexy stuff.
The model is becoming table stakes. The edge is in how deliberately you build the rest.
Stopped chasing every drop, or still testing what lands?
Open-weight models quietly crossed a line this spring.
DeepSeek V4 and Qwen 3.6 Plus are posting 78–80% on SWE-bench Verified — frontier-class agentic coding, open weights, sitting on Hugging Face. The cost-and-control math for building agents just shifted.
But the model stopped being the bottleneck. The harness is the work now: retry logic, real task-completion evals, context that survives long horizons.
Still a jagged frontier — the same model that nails a repo-level refactor can faceplant on something trivial.
That gap is what separates a demo from something you'd actually ship.
What actually made your agents reliable — better evals, tighter scoping, or recovery logic?
@alexabelonix Saw this exact thing on ossaihub. Dropped in bot blockers to test real traffic — time-on-site cratered. Most of that ‘engagement’ was bots.
Pro Blackwell isn't going to drop — the article says it's still climbing. So "now vs wait" is the wrong question. The supply chain's adjusting by routing around the flagship: ~$2K Strix Halo = 128GB unified, used 3090 = 24GB at $1.5K, quantization eats the rest. Real question isn't when to buy the $13K card — it's whether you need it at all.
Installing a third-party agent skill is basically curl | sh into your agent's brain.
It runs with your access — files, keys, network. A "verified author" badge tells you who shipped it, not what it does. And a clean skill can turn malicious on its next update.
~20k places to find MCP skills. Roughly zero that vet what they actually do before they run.
So we built Warden: open-source, local-first, zero-dependency. Pin every skill to a hash (kills rug-pulls), scan it, run it deny-by-default, and flag the ones whose code does more than their manifest admits. Nothing leaves your box. A signal, not a guarantee — and the core is free, forever.
If you run local agents: would you actually pin every skill to a hash and run it deny-by-default? Or is that friction you'd never turn on?
Nobody beat America. America shot itself, and you're cheering at the wrong corpse.
Here's the story you've been fed in the last 48 hours: the US banned its best AI model, China shipped a better one the next day, an open-source model topped the leaderboards, and the export-control hawks just lost the race. Open source wins. Frontier is dead. Unban Fable.
Almost none of that holds up. And the fact that thousands of smart builders are retweeting it anyway is the actual story worth your attention.
Let me walk it.
What actually happened, in order, with receipts
Friday, June 12, 5:21pm ET. The US Commerce Department sent Anthropic an export-control directive suspending all access to Claude Fable 5 and Mythos 5 — for any foreign national, inside or outside the US, including Anthropic's own foreign-born employees. There's no way to screen nationality across millions of users in real time, so Anthropic killed both models globally rather than attempt a partial block. The trigger, per Anthropic's own statement and reporting, was an alleged jailbreak the government considered a national-security risk. White House adviser David Sacks framed it as a reluctant move after Anthropic declined to remediate or de-deploy on the government's timeline.
Saturday, June 13. Zhipu (https://t.co/k2LYQTIXG8) shipped GLM 5.2 to every tier of its GLM Coding Plan. 1M-token context, coding-first, two thinking modes. MIT open weights and a public API promised "next week." The launch messaging leaned hard into "frontier intelligence belongs to everyone" — a direct shot at the ban, with the timing of a press release that was clearly ready to go.
Sunday, June 14. A viral post claims GLM 5.2 is now #1 on a benchmark, beating Fable, at a tenth of the cost and 300 tokens/sec. The AI internet decides open source just won the war.
Three facts, all verifiable. So far so good. Now watch the story fall apart.
The benchmark everyone's citing doesn't say what they think it says
GLM 5.2 launched with zero published benchmarks. Zhipu didn't release numbers. Every serious writeup of the launch says the same thing: vendor claims, third-party verification pending. So where did the viral "#1, 42.8, perfect 100" come from?
One benchmark. Run by the same account that posted the viral thread. That's not independent verification — that's a house publishing its own scorecard and a crowd treating it as gospel.
And the numbers don't even mean what the post implies:
The "42.8 reasoning, #1" headline sits on top of 13.3% raw accuracy on that benchmark's own page. A 42.8 composite built from 13% accuracy is not a frontier triumph. It's a low absolute score on a hard, idiosyncratic test, dressed up as dominance.
The "perfect 100" is on a bullshit-detection category — resistance to fabricating nonsense. Scoring well there is good. It is not "the model is 100% as good," which is how it's being read.
The one number that actually checks out is speed: ~297 tokens/sec measured. Fast. Real. Not the same thing as "best."
So the load-bearing claim of the entire "open source just won" narrative is an unverified, self-published, misread number. If you reposted it, you did Zhipu's marketing for free with a stat you didn't open.
That should bother you more than it does.
The part that kills the "proprietary greed is America's downfall" take
The most popular framing right now is that closed, for-profit models are the US's fatal weakness — that if Fable had been open, this couldn't happen. It's a clean story. It's also backwards.
Fable wasn't pulled by the market. It was pulled by the government, over a safety concern, against the wishes of the company that built it. The for-profit lab is the one fighting to keep it available. The state is the one that switched it off.
And here's the detail that should end the debate: Anthropic stated the same jailbreak likely works on other publicly available models — including OpenAI's GPT-5.5 — which were not subject to the same controls. So the policy didn't even neutralize the capability it was worried about. It removed one American tool from American developers' hands and left the equivalent capability running elsewhere.
That's not the downfall of proprietary AI. That's the downfall of incoherent policy. Different disease. If you misdiagnose it as "closed bad, open good," you'll cheer for the wrong cure.
So what actually beat the US this week? A press release.
Here's the uncomfortable, defensible take: China didn't out-build America in those 48 hours. We don't even have the benchmarks to claim that. What China did was out-narrate America — and it wasn't close.
Zhipu has shipped four flagship-tier coding releases in roughly four months. They IPO'd in Hong Kong in January and raised over half a billion dollars to fund exactly this cadence. They train and serve on domestic chips because export controls forced the muscle to develop. When the US handed them a wound — a banned American flagship — they had "openness belongs to everyone" loaded and fired it the next morning.
The US contribution to this story was to disarm its own developers on a Friday evening over a jailbreak that also affects the model it didn't ban, and then let a self-published leaderboard write the obituary.
This is the lesson the hype crowd and the China-hawk crowd both miss, because it flatters neither: in this cycle, distribution and narrative are beating raw capability. The open-weight coding gap was already closing before any of this — open models hold the majority of tracked leaderboard positions, and the best open coders sit within single digits of the closed frontier on real software-engineering evals. China didn't need GLM 5.2 to be the best model in the world. They needed it to be good enough, available, permissively licensed, and shipped at the exact moment the US looked clumsy. They nailed all four. The fourth one was a gift.
Capability is table stakes now. The race is being run on who can put a competent model in a developer's hands, under a license they trust, with a story they want to repeat. America is currently losing that race not because its models are worse, but because its policy keeps generating the other side's marketing copy.
What this actually means if you build things
Strip the geopolitics and here's the operational reality:
Your best tool can vanish by directive overnight. Fable was generally available on Tuesday and globally dark by Friday. If your agent pipeline had a hard dependency on a single frontier model, you learned this weekend why that's a liability, not a convenience.
Multi-model fallback isn't a nice-to-have anymore; it's the architecture. Open weights you can self-host aren't an ideological preference; they're the only tier of the stack that can't be switched off by someone else's lawyer. That's the genuinely durable argument for open models — not that they topped a leaderboard, but that nobody can remotely revoke them.
And when the next "open model crushes the frontier" post lands — and it will, weekly — open the benchmark before you repost it. Check who ran it. Check the accuracy number under the composite. The thirty seconds you spend verifying is the entire difference between signal and being someone's free distribution.
The models are getting commoditized. Trust isn't. That's the only moat left, and it's the one almost nobody is defending.
The US didn't lose the AI race this week. It lost a narrative skirmish it started itself, and a benchmark nobody bothered to verify is being carved into the headstone.
So here's the real question, and I want builders to actually answer it: when your most capable model can disappear by government order on a Friday night, what does your fallback stack actually look like on Monday morning — and have you ever tested it?
Installing a third-party agent skill is basically curl | sh into your agent's brain.
It runs with your access — files, keys, network. A "verified author" badge tells you who shipped it, not what it does. And a clean skill can turn malicious on its next update.
~20k places to find MCP skills. Roughly zero that vet what they actually do before they run.
So we built Warden: open-source, local-first, zero-dependency. Pin every skill to a hash (kills rug-pulls), scan it, run it deny-by-default, and flag the ones whose code does more than their manifest admits. Nothing leaves your box. A signal, not a guarantee — and the core is free, forever.
If you run local agents: would you actually pin every skill to a hash and run it deny-by-default? Or is that friction you'd never turn on?
Nobody beat America. America shot itself, and you're cheering at the wrong corpse.
Here's the story you've been fed in the last 48 hours: the US banned its best AI model, China shipped a better one the next day, an open-source model topped the leaderboards, and the export-control hawks just lost the race. Open source wins. Frontier is dead. Unban Fable.
Almost none of that holds up. And the fact that thousands of smart builders are retweeting it anyway is the actual story worth your attention.
Let me walk it.
What actually happened, in order, with receipts
Friday, June 12, 5:21pm ET. The US Commerce Department sent Anthropic an export-control directive suspending all access to Claude Fable 5 and Mythos 5 — for any foreign national, inside or outside the US, including Anthropic's own foreign-born employees. There's no way to screen nationality across millions of users in real time, so Anthropic killed both models globally rather than attempt a partial block. The trigger, per Anthropic's own statement and reporting, was an alleged jailbreak the government considered a national-security risk. White House adviser David Sacks framed it as a reluctant move after Anthropic declined to remediate or de-deploy on the government's timeline.
Saturday, June 13. Zhipu (https://t.co/k2LYQTIXG8) shipped GLM 5.2 to every tier of its GLM Coding Plan. 1M-token context, coding-first, two thinking modes. MIT open weights and a public API promised "next week." The launch messaging leaned hard into "frontier intelligence belongs to everyone" — a direct shot at the ban, with the timing of a press release that was clearly ready to go.
Sunday, June 14. A viral post claims GLM 5.2 is now #1 on a benchmark, beating Fable, at a tenth of the cost and 300 tokens/sec. The AI internet decides open source just won the war.
Three facts, all verifiable. So far so good. Now watch the story fall apart.
The benchmark everyone's citing doesn't say what they think it says
GLM 5.2 launched with zero published benchmarks. Zhipu didn't release numbers. Every serious writeup of the launch says the same thing: vendor claims, third-party verification pending. So where did the viral "#1, 42.8, perfect 100" come from?
One benchmark. Run by the same account that posted the viral thread. That's not independent verification — that's a house publishing its own scorecard and a crowd treating it as gospel.
And the numbers don't even mean what the post implies:
The "42.8 reasoning, #1" headline sits on top of 13.3% raw accuracy on that benchmark's own page. A 42.8 composite built from 13% accuracy is not a frontier triumph. It's a low absolute score on a hard, idiosyncratic test, dressed up as dominance.
The "perfect 100" is on a bullshit-detection category — resistance to fabricating nonsense. Scoring well there is good. It is not "the model is 100% as good," which is how it's being read.
The one number that actually checks out is speed: ~297 tokens/sec measured. Fast. Real. Not the same thing as "best."
So the load-bearing claim of the entire "open source just won" narrative is an unverified, self-published, misread number. If you reposted it, you did Zhipu's marketing for free with a stat you didn't open.
That should bother you more than it does.
The part that kills the "proprietary greed is America's downfall" take
The most popular framing right now is that closed, for-profit models are the US's fatal weakness — that if Fable had been open, this couldn't happen. It's a clean story. It's also backwards.
Fable wasn't pulled by the market. It was pulled by the government, over a safety concern, against the wishes of the company that built it. The for-profit lab is the one fighting to keep it available. The state is the one that switched it off.
And here's the detail that should end the debate: Anthropic stated the same jailbreak likely works on other publicly available models — including OpenAI's GPT-5.5 — which were not subject to the same controls. So the policy didn't even neutralize the capability it was worried about. It removed one American tool from American developers' hands and left the equivalent capability running elsewhere.
That's not the downfall of proprietary AI. That's the downfall of incoherent policy. Different disease. If you misdiagnose it as "closed bad, open good," you'll cheer for the wrong cure.
So what actually beat the US this week? A press release.
Here's the uncomfortable, defensible take: China didn't out-build America in those 48 hours. We don't even have the benchmarks to claim that. What China did was out-narrate America — and it wasn't close.
Zhipu has shipped four flagship-tier coding releases in roughly four months. They IPO'd in Hong Kong in January and raised over half a billion dollars to fund exactly this cadence. They train and serve on domestic chips because export controls forced the muscle to develop. When the US handed them a wound — a banned American flagship — they had "openness belongs to everyone" loaded and fired it the next morning.
The US contribution to this story was to disarm its own developers on a Friday evening over a jailbreak that also affects the model it didn't ban, and then let a self-published leaderboard write the obituary.
This is the lesson the hype crowd and the China-hawk crowd both miss, because it flatters neither: in this cycle, distribution and narrative are beating raw capability. The open-weight coding gap was already closing before any of this — open models hold the majority of tracked leaderboard positions, and the best open coders sit within single digits of the closed frontier on real software-engineering evals. China didn't need GLM 5.2 to be the best model in the world. They needed it to be good enough, available, permissively licensed, and shipped at the exact moment the US looked clumsy. They nailed all four. The fourth one was a gift.
Capability is table stakes now. The race is being run on who can put a competent model in a developer's hands, under a license they trust, with a story they want to repeat. America is currently losing that race not because its models are worse, but because its policy keeps generating the other side's marketing copy.
What this actually means if you build things
Strip the geopolitics and here's the operational reality:
Your best tool can vanish by directive overnight. Fable was generally available on Tuesday and globally dark by Friday. If your agent pipeline had a hard dependency on a single frontier model, you learned this weekend why that's a liability, not a convenience.
Multi-model fallback isn't a nice-to-have anymore; it's the architecture. Open weights you can self-host aren't an ideological preference; they're the only tier of the stack that can't be switched off by someone else's lawyer. That's the genuinely durable argument for open models — not that they topped a leaderboard, but that nobody can remotely revoke them.
And when the next "open model crushes the frontier" post lands — and it will, weekly — open the benchmark before you repost it. Check who ran it. Check the accuracy number under the composite. The thirty seconds you spend verifying is the entire difference between signal and being someone's free distribution.
The models are getting commoditized. Trust isn't. That's the only moat left, and it's the one almost nobody is defending.
The US didn't lose the AI race this week. It lost a narrative skirmish it started itself, and a benchmark nobody bothered to verify is being carved into the headstone.
So here's the real question, and I want builders to actually answer it: when your most capable model can disappear by government order on a Friday night, what does your fallback stack actually look like on Monday morning — and have you ever tested it?
Some thought:
1) The Anthropic fiascos is a self-inflicted wound and likely a payback.
2) I spent the night helping many clients off of Anthropic and moving to local open source models and many very large clients will NEVER go back. This has absolutely advantaged open source model from China.
3) This situation NO MATTER WHAT permanently damaged the Anthropic IPO.
4) At 3AM last night I helped a client team move a massive account off of all Anthropic products. This was worth millions of dollars per month and this was the last straw.
5) The fall of Anthropic should not be applauded by anyone. The fall of the company should be viewed as an injury to All US AI COMPANIES.
6) The Anthropic fiasco is not a technical issue, it is a LEADERSHIP issue. If it is not fixed the company is cooked.
7) By the time we end this summer no matter how good Anthropic is, they lose customers, they lose key employees and they ultimately will lose the race.
It was a sad day on top of a massively great day with the SpaceX IPO and one reason I did not post last night.
Dario asked to be regulated, begged to be regulated and yelled to be regulated…
NOW HE IS REGULATED.
You like it now Dario?
We went quiet for a few weeks. Reason: We rebuilt everything. Our own stack, same mission — open-source AI measured, not hyped. V2 is in final gate checks right now: every ranking is math you can verify, every number traces to a source.
What should a directory have to prove before it ranks anything?
Why our sweep runs at the same minute every night: measure at random times and your week-over-week trends quietly lie to you. Fixed window, comparable deltas. Boring discipline, honest charts. How many "trending" lists do you think control for this?