How do LLMs organize emotion concepts?
Excited to share our #ICML2026 paper: we uncover emergent “emotion trees” in LLMs, inspired by emotion wheels from psychology.
📄Emergence of Hierarchical Emotion Organization in Large Language Models
🔗https://t.co/6RxrH9pMiN
What drives in-context learning in LLMs?
New paper: Provable Low-Frequency Bias of In-Context Learning of Representations.
We show LLMs have a low-frequency bias when learning representations in context, offering a theoretical answer to several previously open questions. 🧵👇
🚨 ICML 2025 Paper! 🚨
Excited to announce "Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing."
🔗 https://t.co/If0OHOWSJn
We uncover a new phenomenon, Representation Shattering, to explain why KE edits negatively affect LLMs' reasoning.
🧵👇
🚨 New paper alert!
Linear representation hypothesis (LRH) argues concepts are encoded as **sparse sum of orthogonal directions**, motivating interpretability tools like SAEs. But what if some concepts don’t fit that mold? Would SAEs capture them? 🤔
1/11
New paper---freshly accepted to ICML!
Detailed thread coming soon, but pretty excited about this project. We use synthetic knowledge graphs to study why knowledge editing protocols can screw up model capabilities, finding what we call a "representation shattering" effect!
Q. Why does editing knowledge sometimes reduce overall performance?
A. Representation shatters!
Great insights by a super start undergrad @kento_nishi, supervised expertly by @EkdeepL, with @MayaOkawa, @RahulRam3sh, and Mikail Khona.
New paper–Accepted at #ICLR2025 and also my last PhD paper! 🧑🎓🧵👇
We propose a novel model of how emergent learning curves show up in neural nets’ training by making a connection to the theory of graph percolation!
Paper alert—accepted as a NeurIPS *Spotlight*!🧵👇
We build on our past work relating emergence to task compositionality and analyze the *learning dynamics* of such tasks: we find there exist latent interventions that can elicit them much before input prompting works! 🤯
Excited to share our #NeurIPS2024 spotlight!
https://t.co/XEAnrWo0XP
Key finding: Modern multi-modal models may be hiding far greater capabilities than current protocols can reveal.
Challenge: How do we unlock these "hidden capabilities" through novel extraction protocols?
Paper alert—accepted as a NeurIPS *Spotlight*!🧵👇
We build on our past work relating emergence to task compositionality and analyze the *learning dynamics* of such tasks: we find there exist latent interventions that can elicit them much before input prompting works! 🤯
Our NeurIPS 2023 paper, "Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task," now has a video on YouTube!
It's in Japanese, but English subtitles are available. Check it out: https://t.co/MSm5ZBpFDa
Our NeurIPS 2023 paper, "Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task," now has a video on YouTube!
It's in Japanese, but English subtitles are available. Check it out: https://t.co/MSm5ZBpFDa
More excited than ever to announce $1.7M:
"CBS-NTT Program in Physics of Intelligence at Harvard"! 🧠
With new technology comes new science. The time is ripe to build a better future with "Physics of Intelligence for Trustworthy and Green AI"!
🧵👇https://t.co/60arfZrQAs