we’re moving from “one model does everything”
to “one model that expertly orchestrates the right specialized models”
with long-horizon planning + serious inference
How Anthropic's new results post would read without the PR:
Claude orchestrated open-source protein design models, PXDesign, RFdiffusion, Genie, BoltzGen, from a 30k-token expert prompt and 12,500 H100-hours of compute, and designed binders against 14 of 15 targets. Hit rates of 22–35% against a 10–15% baseline, where some of those tools already report similar numbers on their own.
The orchestration is genuinely impressive. But the open-source models did most of the lifting, and they came from the Baker lab, Columbia, MIT, ByteDance Seed, and most of them were already wet-lab validated before Claude touched them.
Which also sets the ceiling. All these generators share a single PDB-shaped training distribution, so calling four of them doesn't diversify away the blind spot, since they fail together. The targets that worked are the well-studied ones.
So the valid claim is that an agent can now drive this stack competently in the regime where the stack already works.
Instead, we got this announcement:
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
Opus 5 reviews so far:
• argues with instructions
• stops working early
• ignores the handbook
• breaks the existing workflow
• somehow produces the best work
its giving the brilliant coworker everyone hates.
Introducing Claude Code Security, now in limited research preview.
It scans codebases for vulnerabilities and suggests targeted software patches for human review, allowing teams to find and fix issues that traditional tools often miss.
Learn more: https://t.co/n4SZ9EIklG