I wonder if the land of the free has forgotten its principles of upholding freedom, equity and justice and whether the true land of the free is somewhere else!
🚨 Talk Alert!
We’re excited to host Shashata Sawmya (MIT) on May 29 @ 11AM (ICCS 246) for a deep dive into mechanistic interpretability in LLMs — from emergent features to UTune. Don’t miss it! 🧠
#NLP#LLMs#AI#Interpretability#UBC
6️⃣
🌌 Spatial surprise: early-layer concepts vanish mid-stack but re-appear in the final layer—challenging the neat “simple → complex hierarchy” story.
5️⃣
⏰ Temporal emergence:
<3 % concepts fire in the first 1 k steps.
Big surge at 5 k, then two bigger leaps (10 k & 30–40 k).
By 143 k steps > 99 % of concepts are active.
4️⃣
With AUTOINTERP, we label 512 SAE neurons—so each neuron fires for a crisp, human-readable concept. Then we probed for 9 topical concepts — from Physics to Philosophy using a vector database. The framework is called EyeSee.
2️⃣
🤔 Why? LLMs are powerful but opaque. We use sparse auto‑encoders (SAE) to watch interpretable features flicker to life inside Pythia models as they train, grow & traverse layers.
1️⃣
🚀 New paper alert! “The Birth of Knowledge: Emergent Features across Time, Space & Scale in Large Language Models.”
Curious how & when LLMs learn? Thread 👇
#Interpretability#Emergence#LLM#ML
Happy to announce that our work "Wasserstein Distances, Neuronal Entanglement, and Sparsity" with @shashata005, Ilia Markov, @DAlistarh, and @ShavitNir has been accepted into ICLR 2025 as a Spotlight Presentation! Paper link: https://t.co/5aBm3CaZ25
We trained a genomic language model on all observed evolution, which we are calling Evo 2.
The model achieves an unprecedented breadth in capabilities, enabling prediction and design tasks from molecular to genome scale and across all three domains of life.
today we launch deep research, our next agent.
this is like a superpower; experts on demand!
it can go use the internet, do complex research and reasoning, and give you back a report.
it is really good, and can do tasks that would take hours/days and cost hundreds of dollars.
Attending NeurIPS for the first time!
Down to chat about anything from LLM Optimization to Computational Neuroscience (Yep that's the two things I work on :v)