GDM recently found that Gemini's safety eval results are mostly set after SFT, and later stages like RL barely move them. The authors were surprised by this, so we replicated it on Olmo 3 Think. We find the SFT, DPO, and RLVR are within noise of each other on every safety eval we ran.
New post from @secondlookxlab: Rerunning AI safety papers on every frontier release would be pretty easy and valuable 👀
We think that AI safety experiments that are less eval-shaped are rarely rerun on new releases even though there could be lots of actionable information here!
Come work on replications of AI safety research with @zephaniahbroe and I at @SPARexec! @secondlookxlab is hosting 2-4 mentees over the Fall semester.
SPAR runs Sept-Dec, part time, and our application should take <1 hour. You can expect 2+ hours of mentorship from Second Look along with external mentorship from researchers at orgs like Anthropic, X AI, Oxford, Redwood Research, and Geodesic.
https://t.co/TfMjsFeBtT
Fantastic work from Christine!! We replicate single-forward-pass evals on a variety of frontier models and find instances where models have a substantial jump in performance📈
Can frontier models solve hard problems in a single forward pass without thinking? For my first project w/ @secondlookxlab I extended Ryan Greenblatt's no-CoT evals to Fable 5, Opus 5, and GPT-5.6-Sol and saw a substantial jump on some evals!
New replication from Second Look: turns out introspective awareness in LLMs (or at least internal state control) may not be emergent.
great work Finn ‼️‼️
the relationship between superposition and adversarial examples has always been super confusing to me but xijia did a great job making this super digestible in this blog post!!
i'm learning so much from reading the posts that the second look fellows are putting together. highly recommend people read this one!!
We replicate experiments showing the relationship between superposition and adversarial examples in toy models!
Xijia did a great job making this highly confusing topic clear and concise. There is also an open source organized repo which we recommend readers play around with!
Why do adversarial examples exist? Recent research suggests that adversarial attacks leverage interference between features. We replicated superposition in toy models, and found that PGD attacks align with the optimal perturbation direction predicted by latent geometry.
We were able to replicate Jack Lindsey's internal activation control experiments on models as small as 270M parameters!
Because we found the same core result on every model we tested across 3 model families, this is looking like a general property of LLMs🫣
Can LLMs control their internal states? Anthropic found Claude can “think about” a concept on command without saying it. We replicated this in 14 open-weight models and found it to be a general property, maybe not ‘emergent’, appearing in models even as small as 270M parameters.
Can LLMs control their internal states? Anthropic found Claude can “think about” a concept on command without saying it. We replicated this in 14 open-weight models and found it to be a general property, maybe not ‘emergent’, appearing in models even as small as 270M parameters.
In my first project for @secondlookxlab, I replicated and dug deeper into Google Deepmind's "When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors".
shoutout to Arav for producing @secondlookxlab's first replication this Summer!
it's a replication of the DeepMind paper "When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors" (https://t.co/f1miz8oRLH). Check out his work in the LW post!
more replications are on the way 😆
In my first project for @secondlookxlab, I replicated and dug deeper into Google Deepmind's "When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors".
I’m co-founding Second Look Research with @zephaniahbroe and we are accepting summer fellowship applications for 2026!
Fellows will come to the University of Chicago, complete 2-3 replications over 10 weeks (June 15-August 22), and work with external advisors. Fellows will receive a stipend of $10,000 and we will to cover housing and meals.
https://t.co/TfMjsFe3El