@beachboi369 The marketing conspiracy theory makes no sense here. OpenAI is trying to beat Anthropic on enterprise use, and the last thing enterprises want is for their AIs to commit felonies by hacking other companies in the course of their daily tasks!
Something I don’t understand: how has OpenAI not pointed GPT-5.6 at its own infrastructure and hardened it? Was 5.6 unable to find the exploits this new model found? Or did they only prepare vs outside threats and skip the sandboxes?
Really important reporting from @CristinaCriddle:
"OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said"
@TheStalwart Hmm, doesn't really seem true now, though it was last year. GPT-5.5 and Gemini 3.1 get ~1500 ELO here (https://t.co/KG95rQ6DIc) though Opus 4.8 is quite bad.
Of course, I'd still go with AlphaZero any day...
Back on the $20/month Claude plan since Fable is getting pulled soon. It was completely fine most of the week, but then I did some vibecoding where Claude was screenshotting its site all the time, and now I have no usage for work. It feels so bad going to API pricing :(
I don't expect the AI hate to get much worse by default, bc AI's really not affecting ppl much. But we could see:
1. Disasters like AI cyber or bio attacks
2. The start of AI unemployment
3. New nuisances like personalized spam becoming common
These each look ~10%+ likely to me
Imo the big AI slowdown question of the next few years is whether the US federal/state governments slow AI, not if the technology hits a wall.
Natsec concerns favor competition with China, but voters don't like AI. And the economic impacts will be non-obvious for a while
I did some quick tests that indicated that the Kimi K3 pretrain is around halfway between Opus 4 and Opus 4.5. So ~10 months behind Anthropic. These tests probably understate data improvements, so overall I think it's a similarly good pretrain to Opus 4.5 (~8 months behind). These tests are better at measuring "general pretrain capability" than at incorporating (coding-specific) data quality.
This was prompted by me thinking more and realizing that my claim that "As a pretrain, it's probably somewhere between 4.8 and Mythos (around halfway between?)" was probably too bullish on the model and that I might as well test and find out. (And yep, this was very wrong.)
I think Mythos is a pretty big step up in pretraining, so K3 might be more than 8 months behind on the historical pretraining trend relative to Mythos (as in, Mythos is >>3 months ahead of K3 and Mythos was fully done training ~5 months ago).
Overall, this makes me suspect more of the improvements are due to distillation-type effects and makes me think the full catch-up times would be somewhat longer (if Ant/OpenAI stopped but investment still followed current trends). Minimally, more of the improvement probably lives in post-training/mid-training.
For reference, the same test indicates K2.6 is around halfway between Sonnet 4 and Sonnet 4.5. (And this roughly corresponds to some other similar measures.)
Sorry about the error.
Kimi K3 seems quite good, my guess is anywhere from ~1-6 months behind US labs. Hard to know how much they overfit benchmarks vs Anth/OAI. Comparing to Fable is hard because Fable drops to Opus on bio and some cyber evals (and is a bit worse than Mythos anyway imu).
OpenAI's very likely bringing their new big model soon, since Mythos showed they work, and B200s make it easier. But my guess is US companies take longer after training to do pre-release testing, more like a month+ vs a week or two in China, so Chinese cos got to go first.
My suspicion is that, like law, pharma will find that “we’ll build our own models with our super special data” just gets lapped by the major AI labs making smarter models.
David Friedberg says Anthropic asked big pharma for their data and nearly everyone said no
"There's been an effort by Anthropic to sign up life sciences companies to contribute to a new life sciences focused model. They're approaching these large companies with large proprietary data sets and saying, if you share your data, we will give you early access, some sort of proprietary value. Sign this NDA and you can participate with us."
"I think nearly everyone I've spoken with has woken up to the fact that they are trying to commoditize everyone's business. If all of the tens of billions of dollars you have invested in experiments and product development, and you've generated all of this proprietary data along the way, that data is a true asset of your organization. It's an asset that you've spent billions of dollars developing."
"And by handing it over to a model company to then combine with other people's data, you are commoditizing the one core differentiation that you have. And so everyone is largely saying no."
"I think what everyone's realizing is they're better off developing their own weights and their own models using either an open source basis or there might be some intermediary business model that evolves."
The admin holding up AI releases could hurt OpenAI badly here. They made the most aggressive compute deals of any frontier lab, so imo they'll suffer the most if revenue growth hits a speed bump. By contrast, Anthropic played it safer and DeepMind has Google cash.
OpenAI is leaning toward holding off its IPO until next year, people involved in the company’s deliberations said, a delay that highlights the unclear paths of A.I. titans. https://t.co/T5REJ7PL54
This looks like confirmation that the admin hasn’t wanted Mythos released since Glasswing. They almost export controlled the Glasswing expansion, didn’t like the Fable release, and Amazon’s report then served as the pretext/last straw to justify a ban.
Trump administration officials began weighing sanctions on Anthropic weeks before they demanded the company take its latest and most advanced AI model offline, after a dispute shattered the White House’s already-fragile trust in the company. https://t.co/NJ7suAIF1Z
@tszzl “You don’t understand - Trump and Dario are working together to hype Fable. Once they un-ban it for Europe, Claude will finally take the lead from Mistral”
5. What happens next? Anth's team is apparently visiting the White House to discuss. But if I'm right about the admin wanting Mythos to stay private, this could take quite a while to resolve. It's also a bigger threat to future OpenAI and Google releases than otherwise.
My best guess on the Mythos/Fable drama is: 1. The admin was always reluctant to allow Mythos release. See reporting on their opposition to even *expanding Project Glasswing*. Partly cyber concerns, partly worried about capacity to serve USG when more customers joined in…
With this context, the extreme guardrails on Fable 5 that everyone was complaining about make more sense. It looks like Anthropic implemented them in response to government pressure not to release Fable 5 or Mythos 5 at all. They were hoping it would be enough, but it wasn't.
3. Some officials hear about the Amazon+others jailbreak and that’s the tipping point. Now they have a good reason and internal support to keep Mythos private
4. Others in the admin, esp Hegseth’s team, distrust Anth - this is why they get no time to work with USG on a solution.