I gave a lightning talk at the mechanistic interpretability workshop at ICML this year. Thanks to the organizers, who encouraged us to share our high-level takes; my talk ended up being quite high-level and personal indeed. Blog post version:
https://t.co/F0QCXvh655
We’re in the process of scaling up parameter decomposition from 67M models to 8B. But there are already tonnes of interesting threads like this that you can explore that will be evergreen.
We removed an LM's ability to speak German by fine-tuning on only 4 German tokens.
As part of a 1-day hackathon with our product Silico, we removed a 67M-parameter language model's ability to predict German text, by tuning only a scalar factor on one subcomponent of the weights. (1/6)
My team at @GoodfireAI has been cooking up a new way to do interpretability: decompose a language model’s weights, not its activations.
Our decomposition natively handles attention (!) and behaves less like a lookup table and more like a generalizing algorithm. (1/6)
The most popular way to interpret AI is missing the bigger picture.
Models think in curved shapes. But sparse autoencoders (SAEs) work with straight lines.
Can they still capture models’ curved neural geometry? Yes, but not how you might think! (1/7)
Neural networks might speak English, but they think in shapes.
Understanding their rich *neural geometry* is key to understanding how they work – and to debugging and controlling them with precision.
Starting today, we’re releasing a series of posts on this research agenda. 🧵
Coding agents should be treated as untrusted by default
We built Watcher as an MDM/EDR for coding agents. Security teams set hard boundaries (blocked commands, req. monitors), engineers set the rest
Watcher will support guardrails & monitors for Claude Code, Codex, Cursor, etc
I’m obviously biased, but I think this method (or a close neighbour) can give us the interpretable, mechanistic faithful components of an arbitrary neural network that we’ve long been looking for.
(Sound on!)
My team at @GoodfireAI has been cooking up a new way to do interpretability: decompose a language model’s weights, not its activations.
Our decomposition natively handles attention (!) and behaves less like a lookup table and more like a generalizing algorithm. (1/6)
I'm extremely grateful to have worked on this project. It's the closest thing I've seen to the vision that inspired me when I originally got into Mech Interp. I think we truly can and will deeply understand language models!
We achieved state-of-the-art performance in predicting which of 4.2 million genetic variants cause diseases by interpreting a genomics model, in a new preprint with @MayoClinic.
We're now releasing an open source database for all variants in the NIH's clinvar database. 🧵(1/8)
New Paper! RL can teach our models to solve math or code, but open-ended tasks — which make verification expensive or even impossible — remain difficult to optimize. LLMs-as-Judges help, but often struggle to retrieve information even when it is present.
Reinforcement Learning from Feature Rewards (RLFR) provides a solution. Extracting model beliefs via interpretability reveals a well-calibrated reward signal that permits scalable training.
We raised a $150M Series B at a $1.25B valuation to fundamentally change the field of AI. Scaling is powerful, but we can't intentionally design what we don't understand.
This could be the biggest result to date in the application of interpretability to science (though maybe there are some neuroscience results I’m overlooking?).
Pretty damn cool to work alongside people doing this.
We've identified a novel class of biomarkers for Alzheimer's detection - using interpretability - with @PrimaMente.
How we did it, and how interpretability can power scientific discovery in the age of digital biology: (1/6)
🔶Introducing our new method, SPD🔶
I’m now very excited about scaling this thing and testing it on many different architectures ( https://t.co/mjBvuC66lv if you want to try!). I think this direction might give us the faithful network decompositions we’ve always hoped for.
A few months ago, we published
Attribution-based parameter decomposition -- a method for decomposing a network's parameters for interpretability.
But it was janky and didn't scale.
Today, we published a new, better algorithm called
🔶Stochastic Parameter Decomposition!🔶
I've joined @GoodfireAI (London team) because I think it's the best place to develop and scale fundamental interpretability techniques.
Doing this well requires compute, ambition, and most of all, great people. Goodfire has all of these.
This means I'm saying goodbye to @apolloresearch. I'm continually impressed with how motivated and unwavering in their ideals the folks at Apollo are. We've had a fun and successful couple of years together and I'll continue to be their biggest fan, just from the sidelines now.