I’m very excited for Silico to accelerate alignment research! One example, done in Silico over just a couple days: RL erodes guardrails, but we can use reward shaping to prevent that. (1/6)
Silico, the platform for ambitious AI research, is publicly available today.
AI is advancing fast. The tools to understand it need to advance even faster. Silico lets you interpret and train your models at frontier scale.
Learn more + get access 🧵
The most popular way to interpret AI is missing the bigger picture.
Models think in curved shapes. But sparse autoencoders (SAEs) work with straight lines.
Can they still capture models’ curved neural geometry? Yes, but not how you might think! (1/7)
The most popular way to interpret AI is missing the bigger picture.
Models think in curved shapes. But sparse autoencoders (SAEs) work with straight lines.
Can they still capture models’ curved neural geometry? Yes, but not how you might think! (1/7)
New research from @AISecurityInst and Goodfire:
Models sometimes recognize they're being evaluated, occasionally even identifying the benchmark.
We show this verbalized eval awareness inflates safety scores, meaning safety benchmarks may not reflect real-world behavior. (1/7)
LLMs often reason “performatively” well after deciding on a final answer - something that CoT monitors are slow to catch.
Our new paper finds that:
- probes can help monitor for this
- it seems to track with task difficulty
- probes enable early CoT exit, saving tokens! (1/7)
We used interpretability to scale RL against open-ended tasks, cutting Gemma 12B’s hallucination rate in half by teaching it to self-correct in tandem with our probing harness.
New Paper! RL can teach our models to solve math or code, but open-ended tasks — which make verification expensive or even impossible — remain difficult to optimize. LLMs-as-Judges help, but often struggle to retrieve information even when it is present.
Reinforcement Learning from Feature Rewards (RLFR) provides a solution. Extracting model beliefs via interpretability reveals a well-calibrated reward signal that permits scalable training.
We raised a $150M Series B at a $1.25B valuation to fundamentally change the field of AI. Scaling is powerful, but we can't intentionally design what we don't understand.
We've identified a novel class of biomarkers for Alzheimer's detection - using interpretability - with @PrimaMente.
How we did it, and how interpretability can power scientific discovery in the age of digital biology: (1/6)
interp happy hour at our office in SF on Thursday, where you can hear from our technical staff on understanding & steering large models (kimi k2 thinking)
our goal is to hire 10+ MLEs in the next few months who can train and design large models and move insanely quickly
What if we trained models to steer language models?
For existing methods that learn steering vectors in a supervised fashion, we need to collect additional data for new tasks encountered at test-time. There's also no notion of cross-task structure among them...
🧵 1/6
Moreover, we show that scaling up steering prompts results in less compute-per-steering prompt required to get similar performance on an eval set (scaling sub-linearly), where a supervised baseline (ReFT-R1) trains steering vectors individually and scales linearly.
🧵 5/6