Most interpretability starts after the model is trained.
We tried the opposite.
Build interpretability into the model during training, so its internal computation can later be inspected and steered.
𝑺𝒕𝒆𝒆𝒓𝒍𝒊𝒏𝒈 is our attempt to see how far this idea can scale.
Full technical report on arxiv:
https://t.co/rBtLgvYNUB
We can now edit a cell that won't respond to a drug into one that does, in silico.
With the BRAID team at @genentech, we built scCBGM: an interpretable #AI model that extracts the mechanism behind a cell's response, then lets you act on it. 🧵
technical deep dive on a new dataloader that allows you to train models with explanations. It allows you to annotate token spans in a token stream with any metadata you want. Incredible effort by a *single* member of our team.
Working at a startup, you get to cover many surprising fields, such as inventing a new dataloader. I used both my experience from my Node.js core developer days (2012-2019), and my recent experience as an ML researcher. Read the blogspot to understand all the technical details.
NEW: We built a new dataloader that streams metadata alongside the token stream to train concept-aware, inherently interpretable language models at scale. That’s how we’re able to see how Steerling-8B reasons & steer.
Breakdown ↓↓↓
Today we’re announcing a finding that breaks a core assumption in AI: that bigger models are harder to understand.
We show the opposite. When interpretability is built into training, models become MORE understandable as they become more capable.
Here is #Clarity: an interpretable chatbot where you can inspect why the model produced an output.
Concepts. Training data. Control. All in one interface.
Personally, I’m most excited about what this could mean for safe AI: moving beyond giant models optimized only for downstream accuracy, toward models we can actually understand, trust, and control.
Say hello to Clarity 👋 An interpretable AI platform that actually lets you see -- and control -- what’s under the hood.
1. Full concept-level transparency
2. Direct traceability to training data
3. The best part? You can steer model responses in real-time.
At Guide Labs, we are building LLMs that help humans, with interpretability engineered from the inside out.
We believe interpretability is how we make AI reliable, trusthworthy, and capable of driving real-world progress.
Read more about our vision: https://t.co/CvrMMSWjzG
Play with opensource version of our model on GitHub: https://t.co/PIVwJgleFP
#ResponsibleAI #GuideLabs
UPDATE: We’re open sourcing a token-level attribution system that breaks down every Steerling-8b token generated into:
- what we explicitly taught it
- what it learned on its own
- what we still can’t explain
Our interpretability seminar today (22/04) features @asalam_91 from @guidelabsai presenting the largest interpretable model in the world!
🗓️ Topic
Steerling-8B: Inherently Interpretable Language Model
Tune in live at 5:30 CEST / 8:30 PST: 📷https://t.co/yLa75UVt6J
14.8 million documents. 95 million chunks. 10 billion tokens. 16.8K human-understandable concepts.
Interpretability comes from datasets with structure.