AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public.
Today, we're sharing three measurements that help track AI development:
1. How much AI R&D is done by AI.
2. How well AI agents are overseen.
3. How compute is allocated.
We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them.
As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information.
Read the full post and methodology: https://t.co/iPFz8Z4ugE
Claude Code 这次更新,改的是人和 AI 的分工。
官方演示里,程序员往对话框连甩五句碎话,错字都没改。Claude 自己认出五个意图,合并成三件事,开三条线程同时干。结账重复扣款和 Stripe 沙盒测试,它判断是同一件事,放进同一条线程。
中途插一句"少一点可能更好",话会自动转进对应线程,没过多久多出一版方案,选完直接开 PR。
每条线程都是云端一个完整会话,合上电脑照样跑。代价是额度烧得快。
目前只给部分 Pro 和 Max 用户测试。
Projects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop.
In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.