This is a good teaser, but everyone's just waiting for the "caterpillar plot."
In seriousness, I've long described this as the most influential NAEP plot, despite my quibbles (see Ho & Haertel). As much as I care about trends, I know many others care about standards. @NAEP_NCES
Dear international colleagues,
I’ll be on sabbatical in 2026–27 and would be excited to give seminars at universities abroad.
If my work on teacher labor markets, tutoring, effect sizes, time in school, or climate impacts on education might be of interest, please reach out.
Today we're launching Claude for Teachers -- premium @claudeai and Cowork, free for every US teacher.
Teachers have been experimenting with AI for a while. But they told us they wanted something curriculum-aligned, evidence-based, and able to work in the background while they focus on their students.
Four things that I think make this special:
Do I know Any Public School leaders on X?
Next year we'll be giving out ~$10m of free expert-tutoring services and I want to connect with public school leaders who are interested in partnering with us.
This is genuinely very cool. Science infrastructure is extremely clunky.
Doing my PhD, I spent so much time doing extremely boring and annoying things like switching between file formats that didn't work together and so on. I think it would be super cool if these tools reduced the time spent on boring stuff and allowed scientists to focus more on science.
Some news: As of June 30, I'll be on leave from Stanford at Anthropic. I'm joining the Anthropic Institute, where I'll continue my research on AI and our economic future and give seminars and talks as always. 1/3
🚨Our tutoring meta-analysis is now online at RER.
@BethSchueler, Grace Falken & I analyzed 263 RCTs.
What we learned has important implications for the future of tutoring and the use of meta-analyses to inform policy.
Open access:
https://t.co/G0Mf3MGIs3
🧵
Fable 5 is doing something wild on our FrogsGame post-training task.
It trains a weaker model to solve the puzzle, peaks at 68%, and produces the only ~10x improvement we see across the benchmark.
It spent 17 hours, 25M tokens without human in sight. 34% pass@1, while every other frontier model averages under 4%.
We will publish a more detailed analysis soon.
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use.
Its capabilities exceed those of any model we’ve ever made generally available.
This morning I opened the laptop to check email, realized I needed some additional (simple) robustness checks for a paper, opened Claude and entered my prompt; then my toddler climbed on my lap + I cuddled himwhile answering simple Qs so the agent could do tedious coding