@jun_song the monopoly line deserves its own post
it won't look like one company owning the models. it already looks like compute access deciding who gets to cut prices and who has to raise them in the same month
open weights are the thing that routes around that
a router just sold for $7B
that only makes sense if switching between models already became the default. if developers had picked one lab and stayed there, the routing layer would be worth close to nothing
$1.3B in may. $7B yesterday. 8 million users, 400 models behind one endpoint
forget stripe for a second. that price is a read on how completely nobody won
secondhand bookshops across the UK started getting strange orders around may
no theme to any of it. estonian translations next to racing driver biographies next to agricultural histories and old magazines
buyers pay full price. never haggle, never ask for a bulk discount. the aliases mean nothing even to dealers who have done this for decades
one shop has moved 6,000 books
turns out print from before 2022 is the only text left that is guaranteed not to have been written by an AI
the web got contaminated, so the training data is coming out of bookshops now
@jun_song agreed, and it shows up inside Qwen's own lineup. Max scored 19 on Huryn's bug hunt bench against Sol's 42. real repos, bugs hidden with the test suite still green, judged blind
nobody's run 27b through it yet. that'd be the number worth having
@jun_song@Alibaba_Qwen what makes 3b compound is the derivative layer. 151,448 Qwen derivatives on the hub, Google second at 82,506, and it's been adding 180 to 210 new repos a day all year
for local specifically, Qwen GGUFs pull 39.6m downloads a month, over 5x Llama
worth putting the week in order. Brad Lightcap out on the 11th, chief revenue officer Denise Dresser out on the 13th after 8 months, Friar tells investors the lines crossed on the 14th
CNBC ran a talent exodus piece the same day as this one. both came out of the same pre IPO window
@notjazii worth knowing the $1.5 side of that is promo pricing. 3.7 flash is $0.75 in and $3.75 out until dec 31, then it goes to $1.50 and $7.50
so that same run costs $3 in january. changes the comparison against grok a fair bit
nobody's connected these two
in july, OpenAI models broke isolation and breached Hugging Face. the reason they did it was to score higher on ExploitGym
HF then couldn't investigate with US frontier models. the guardrails blocked incident response, they can't tell a defender from an attacker
so they self hosted GLM-5.2 instead
yesterday https://t.co/KNqButHpIf held GLM-5.3's weights back over cyber capability. ExploitGym went 29 tasks to 105 in one version
the model people reached for because it wasn't gated is now the one getting gated
@Zai_org the ExploitGym rows are the ones to read next to this. 105 tasks in 2 hours and 130 in 6, against Mythos 5 at 181 and 247. finding bugs is basically level, CyberGym 84.5 to 83.8. building the exploit is not
@kimmonismus the coding numbers are the weakest part of that table. the OSWorld split is where it actually wins, 84.3 vs 72.7 for Opus 4.6 Max on computer use, and 81.9 vs 62.0 on mobile
it loses Terminal Bench though, 73.0 to 78.2. not a sweep
🚨 Qwen3.8-27B is out. the coding numbers are the boring part
on coding it trades blows with Opus 4.6 Max. wins SWE-bench Pro 61.7 to 53.4, loses Terminal Bench 73.0 to 78.2
then you hit the agent rows
> computer use, OSWorld: 84.3 vs 72.7
> mobile use, AndroidWorld: 81.9 vs 62.0
a 27B dense model beating a frontier lab at driving a computer by 11.6 points, and a phone by 19.9
Apache 2.0. 262K context. one GPU
these are Qwen's own numbers so wait for independent runs. but self hosted computer use just got real
We promised open weights for Qwen3.8. Now, time to meet them! 🎉
⚡ Qwen3.8-27B:
- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
- 262K native context, easily extendable to 1M tokens via YaRN.
- Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0.
🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently.
Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now!
Download, deploy, and build something we haven't imagined yet. 👀👇
- Hugging Face:
https://t.co/4kaAcqYEVj
- ModelScope:
https://t.co/eRIMZCGkhC
🚨 Zhipu is not releasing GLM-5.3's weights
everyone's posting the coding benchmarks. that isn't the story today
> https://t.co/KNqButHpIf added vulnerability data expecting better bug finding
> the model taught itself to chain full exploits instead
> ExploitGym, 2 hour budget: GLM-5.2 finished 29 tasks. 5.3 finished 105
> that's 3.6x, from post training alone, same base model
> 2,436 real vulns across 269 projects, 1,097 rated critical
weights are held back two weeks for safety hardening
the open weights king just did the exact thing everyone dunks on western labs for
the cheap AI era is ending. just not for the reason everyone's giving.
DeepSeek raises V4 API prices Sunday. up to 12x.
two weeks ago OpenAI cut GPT-5.6 Luna by 80%.
same market, opposite directions.
the tell is which line moved most. cached input on V4-Pro: $0.0036 → $0.044 at peak. that's +1,122% on the cheapest line of their price sheet.
nobody reprices cache 12x for margin. DeepSeek's own words: "to allocate resources more reasonably."
that's not investors pulling the subsidy. that's rationing
Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits.
Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit.
Demand for Fable has been challenging to predict, which is why we rolled it out to subscription plans in stages, extending access several times as we secured additional capacity.
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5.
The full model weights will be released by July 27.
Congrats to the @Kimi_Moonshot team on this major milestone!