I’m co-founding Second Look Research with @zephaniahbroe and we are accepting summer fellowship applications for 2026!
Fellows will come to the University of Chicago, complete 2-3 replications over 10 weeks (June 15-August 22), and work with external advisors. Fellows will receive a stipend of $10,000 and we will to cover housing and meals.
https://t.co/TfMjsFe3El
The case for replications is somewhat self-evident. Yet, across scientific disciplines, replication is systematically neglected because the work is unsexy and novel research is exciting.
Someone needs to do the boring work!
New blog post from me and @Yixiong_Hao: "Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced"
more thoughts in thread🧵
https://t.co/Gn06UHSdDl
By replicating and stress-testing, we both:
1. verify that the experimental result can be reproduced
2. empower the rest of the AI safety community to build on the research the lab started
David Sachs in a recent interview says:
“If Dario believes that we are going to have the worst outcomes my question for him is, why are you doing this?”
Lol unintentionally based take. Anthropic should pause.
Unfortunately, it looks like findings on Astra's no-CoT capabilities replicate.
Astra can compute more in a single forward pass than its predecessors raising concerns about how much we can expect CoT monitoring to hold for it or for future models.
In case you need more confirmation that GPT-Astra is terrifyingly good at no-CoT computation: I ran my replication of Ryan Greenblatt’s no-CoT evals on Astra (+ Gemini 3.1 Pro, Kimi k3, Fable 5.1). It is a qualitative jump on every dataset, especially competition math and multi-hop reasoning. 🧵
New LessWrong post by @secondlookxlab: SFT Also Drives Safety Eval Results in Olmo 3
Prior GDM work found safety eval results remain mostly fixed after SFT. This post finds the same in Olmo 3. I think the GDM results deserve more discussion and this replication is a good start!
GDM recently found that Gemini's safety eval results are mostly set after SFT, and later stages like RL barely move them. The authors were surprised by this, so we replicated it on Olmo 3 Think. We find the SFT, DPO, and RLVR are within noise of each other on every safety eval we ran.
XLab AI Safety Group's Nolan Johnson was interviewed on FOX 32 Chicago yesterday!
Nolan talked about recent shakeups at Anthropic, the case for AI x-risk, and the UChicago AI safety group. Full segment linked below.
XLab AI Safety Group's Nolan Johnson was interviewed on FOX 32 Chicago yesterday!
Nolan talked about recent shakeups at Anthropic, the case for AI x-risk, and the UChicago AI safety group. Full segment linked below.
I'm sort of surprised how much attention this is getting considering that Evan has been talking about this for years! In 2020 (before public release of Chat GPT!): "I sort of see the problem that we're trying to solve as the problem of AI existential risk."
https://t.co/S9RdPcPc64
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
@ishanjmukherjee Fair enough! My point is specifically about the wording and framing. I think it’s pretty important to point out when labs/anyone conflate concepts relating to alignment.
This is extremely bad. I would have thought that the almost universal consensus of this being a terrible idea would dissuade OAI from it, but alas, you can never quite trust them with anything. This is horrible news for safety and for society as a whole.
good o’ days talking about inner/outer alignment in a dim staircase
“Many of the building's walls were concrete. If you are a sufficiently nerdy person, you would know this is great news because you can write on concrete with chalk, so everything vertical becomes a blackboard.”