We’re optimistic that every training run can be monitored and that the current rates of reward hacking may soon be a thing of the past.
Read the full post + paper: https://t.co/aMB75AbyaX
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied!
We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
embedded evaluations are a great idea, and the team at @GoodfireAI would love to help with white box evaluations
the goal of white box evaluations could be to predict future misaligned behavior directly from the internal mechanisms of models, as well as stress test activation monitors
Models usually know when they’re doing something wrong.
Probes let us “read their minds” to detect offensive cyber intent, reward hacking, strategic deception, and more. If you build or serve models, they’re your first line of defense.
Our latest post explains how they work & how to use them - the first of a series of educational posts on applied AI interpretability. 🧵
"Why am I optimistic about interpretability? It's partly because I think we're starting to have really good traction - but also because I can imagine this incredible speedup."
Catch our Chief Scientist @banburismus_ on the latest episode of @MLStreetTalk!
in light of multiple models breaking containment, we've decided to focus our research at @GoodfireAI to solving AI alignment via interpretability.
the hugging face incident is a turning point for the world where AI safety gets real. i am personally very concerned. i'm glad that the AI labs are taking these issues seriously, but i think it's unlikely that we can fully align AI models without significant progress in interp.
taking a big swing at solving interpretability to align models has always been our long term vision, but we have concrete problems that we can tackle with increased urgency.
our research will focus on two areas - interp foundations and safety applications. we want to take ambitious swings to fully reverse engineer neural networks and steer them during training. we believe this will be deeply important in the long term. we'll also turn more attention to immediate safety applications, starting with identifying and preventing reward hacking, cybersecurity risk, and biorisk directly from model internals
much more to come here. if you'd like to directly work on solving AI alignment via interpretability, consider joining us
With just one prompt, we taught an LLM to see - then looked at its representations to debug where it fell short.
Qwen 3 8B can read text, but has no way to see or understand images. Silico trained a vision adapter that matches the official Qwen 3 VL 8B in multiple benchmarks. 🧵
Silico solved 2 open cases of the approximate counting colorings conjecture. @fran_perni described its proof style as an “experimental scientist at heart doing math.”
"I think each of these would make a nice paper at a top theoretical computer science conference."
Every human has millions of mutations in their genes. How do we know which ones are harmful?
@jaanakprashar built MAPS, a Mechanistic Atlas of Protein Sequences, explaining 2.1 million genetic variants — and asking not just *whether* a mutation is harmful, but *why* 🧵
Silico, the platform for ambitious AI research, is publicly available today.
AI is advancing fast. The tools to understand it need to advance even faster. Silico lets you interpret and train your models at frontier scale.
Learn more + get access 🧵
You ask an AI model a question. Why does it answer the way it does?
Using the model's own neurons, we can trace how it makes decisions — and steer it toward better ones.
Case study: we tested an LLM and found that it sometimes endorses drunk driving.🧵 (1/5)
In a recent hackathon, we trained an LLM to label its own neurons. But our RL run collapsed. 94% of labels opened with the word "texts."
Instead of reworking our data + retraining, we edited the weights directly - dropping “texts” from 94% to 5% with minimal side effects. (1/5)
There's a "bouba-kiki space" inside LLMs
"Spiky-sounding" and "round-sounding" words align on opposite ends of a specific direction in Llama & Gemma's activations, independent of their meaning, mirroring the famous bouba-kiki effect
PCA reveals beautiful geometry inside a protein language model. But does the model actually use it?
I explored how protein folds—like this beta-propeller fold—are represented in ESMC-6B using the interpretability tools built into Silico, @GoodfireAI's research platform 🧵
3 months ago, we used interpretability to predict which of 4.2 million genetic variants cause disease.
Now, we've validated several of those predictions with real-world datasets, including a national biobank, clinical data, and RNA sequencing data.
4 examples: (1/6)
> replicate J-space on GLM 5.2
> train a reward model and run RL to reduce hallucinations
> show me how this model makes cancer predictions
Using our platform Silico is like having a team of AI researchers ready to run experiments like these.
Private beta is open now. 🧵 (1/6)
Can LLMs predict the next World Cup champion?
Goodfire partnered with @EternisAI to improve how LLM forecasters use available evidence and manage uncertainty.
We found models were overconfident in their predictions – but probes significantly improved calibration. (1/6)
If models think in shapes, our tools should too.
Our latest research: Block-Sparse Featurizers (BSFs), a new way to find concepts in model activations - using multidimensional “blocks” instead of single directions. (1/9)
We removed an LM's ability to speak German by fine-tuning on only 4 German tokens.
As part of a 1-day hackathon with our product Silico, we removed a 67M-parameter language model's ability to predict German text, by tuning only a scalar factor on one subcomponent of the weights. (1/6)
Stories have shapes: a comedy rises toward joy; a tragedy falls into loss.
Inside an LLM, that’s visible more literally: as an LLM reads a story, its internal activations trace a wandering path that reflects the model’s sense of what kind of story it is reading. (1/5)