Magnífico: en una muy patriótica inauguración portuaria en Estados Unidos se levanta un poco de viento, vuela una bandera estratégicamente colocada y deja al descubierto la empresa china que ha construido las grúas.
Occidente cada vez fabrica menos y depende más.
Even @OpenAI's recent Erdős breakthrough didn't convince me that LLMs can do general math research. This changed my mind..
Using a clever 'prover-verifier' LLM loop, this harness solved 9 substantial open problems in Theoretical CS, including one that kept me up at night for 2 years.
Incredible work by my former Columbia collaborator @binghuip, @runzhou_tao, Steven Wang & @HantaoYu_Theory.
The plan is to expand this to ALL fields of science. Stay tuned.
It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know what risks everyone will face if/when open source reaches Mythos class & (b) whether they are doing enough or too much to prevent those risks.
There’s a big misconception about how GLM 5.2 was trained. Yes, they distilled Claude and GPT 5.5 — but distillation is not how they matched Opus quality. Distillation only fixed the cold start problem in RL.
RLing an agentic coding model isn’t rocket science. In simplified terms:
1. RL needs trajectories — rollouts where the model actually completed a task in some env
2. No successful trajectory on a task = zero gradient = you can’t RL it. This is the cold start problem
3. Distillation solves it. You seed your model with knowledge from a smarter one (Claude, GPT) on tasks it can’t do yet
4. Now it produces positive trajectories on those tasks
5. RL on those trajectories and hill climb agentic coding
6. At that point you no longer need to distill and can solely hill climb RL to better models
This is an interesting curve. I’d argue it’s harder to get to Opus 4.8 from scratch than to go from Opus 4.8 → Fable/Mythos tier.
GLM 5.2 is already producing positive trajectories, so they have plenty to RL on — they’ll keep climbing to Mythos quality without distilling any further. They no longer need American models.
Cool way to use Claude Code: deciphering Linear A, a 3500 year old written language from Crete
https://t.co/Aqd4ZG7Cum
Hope this holds up in peer review! 🤞
Two things are true:
(1) Anthropic (or parts of it) are absolutely and sincerely worried about the misuse of Mythos-class models & have put in excessive safeguards until they are confident it will not be misused
(2) They have not succeeded in explaining/convincing people of this
@miriamgonp@AnthropicAI Está limitando todas las consultas que se le hacen sobre medicina, biología y demás… así que de momento me da que de poco sirve ahí.
@DeryaTR_@iScienceLuvr Parece que cualquier petición que tenga que ver con algo relacionado con salud lo bloquea por defecto. Quería analizar unos datos producidos por mí y también se niega a hacerlo. ¿Para qué publican benchmarks de salud si luego no dejan a los usuarios probarlo?