> replicate J-space on GLM 5.2
> train a reward model and run RL to reduce hallucinations
> show me how this model makes cancer predictions
Using our platform Silico is like having a team of AI researchers ready to run experiments like these.
Private beta is open now. 🧵 (1/6)
Stories have shapes: a comedy rises toward joy; a tragedy falls into loss.
Inside an LLM, that’s visible more literally: as an LLM reads a story, its internal activations trace a wandering path that reflects the model’s sense of what kind of story it is reading. (1/5)
The most popular way to interpret AI is missing the bigger picture.
Models think in curved shapes. But sparse autoencoders (SAEs) work with straight lines.
Can they still capture models’ curved neural geometry? Yes, but not how you might think! (1/7)
Neural networks might speak English, but they think in shapes.
Understanding their rich *neural geometry* is key to understanding how they work – and to debugging and controlling them with precision.
Starting today, we’re releasing a series of posts on this research agenda. 🧵
My team at @GoodfireAI has been cooking up a new way to do interpretability: decompose a language model’s weights, not its activations.
Our decomposition natively handles attention (!) and behaves less like a lookup table and more like a generalizing algorithm. (1/6)
Introducing Silico: the platform for building AI models with the precision of written software.
Silico lets researchers and engineers see inside their models, debug failures, and intentionally design them from the ground up.
Early access is open now. 🧵(1/10)
We achieved state-of-the-art performance in predicting which of 4.2 million genetic variants cause diseases by interpreting a genomics model, in a new preprint with @MayoClinic.
We're now releasing an open source database for all variants in the NIH's clinvar database. 🧵(1/8)
New Paper! RL can teach our models to solve math or code, but open-ended tasks — which make verification expensive or even impossible — remain difficult to optimize. LLMs-as-Judges help, but often struggle to retrieve information even when it is present.
Reinforcement Learning from Feature Rewards (RLFR) provides a solution. Extracting model beliefs via interpretability reveals a well-calibrated reward signal that permits scalable training.
Every engineering discipline has been gated by fundamental science (think steam engines before thermodynamics), and AI is at that inflection point now.
We raised a $150M Series B at a $1.25B valuation to fundamentally change the field of AI. Scaling is powerful, but we can't intentionally design what we don't understand.
We raised a $150M Series B at a $1.25B valuation to fundamentally change the field of AI. Scaling is powerful, but we can't intentionally design what we don't understand.
Incredibly proud of what the team has achieved here! This is just the very beginning of realizing our vision for interpretability as an engine for novel scientific discovery. Excited for what's to come
We've identified a novel class of biomarkers for Alzheimer's detection - using interpretability - with @PrimaMente.
How we did it, and how interpretability can power scientific discovery in the age of digital biology: (1/6)
Every single one of these $100M+ companies were started by alumni from a single Computer Science club in a non-American high school.
Cartesia
Inception Labs
General Catalyst CVF
Wispr Flow
Affinity
Snapdeal
Sugar
boAt
It's Exun Clan in Delhi Public School, RK Puram in India.
Why use LLM-as-a-judge when you can get the same performance for 15–500x cheaper?
Our new research with @RakutenGroup on PII detection finds that SAE probes:
- transfer from synthetic to real data better than normal probes
- match GPT-5 Mini performance at 1/15 the cost
(1/6)
Agents for experimental research != agents for software development.
This is a key lesson we've learned after several months refining agentic workflows!
More takeaways on effectively using experimenter agents + a key tool we're open-sourcing to enable them: 🧵
Startups are an emotional journey with high highs and low lows. Be stoked for your friends who are having an easy time - but when they are going through the low lows, be there for them in every way you can. They will appreciate it.
The new GPT-4o image model is autoregressive, not diffusion. What’s the difference?
Autoregressive: imgs emerge sequentially, like Miyazaki carefully drawing each animation cell
Diffusion: imgs crystallize slowly from noise, like watching a Ghibli watercolor bloom into clarity
3/ Details are less important than the mythology.
This is so true in early stage pitching.
The mythology you want to build and deliver is fundamentally what an investor underwrites. The details are simply useful in making the mythology seem within reach.