π§ To reason over text and track entities, we find that language models use three types of 'pointers'!
They were thought to rely only on a positional oneβbut when many entities appear, that system breaks down.
Our new paper shows what these pointers are and how they interact π
Can we precisely erase conceptual knowledge from LLM parameters?
Most methods are shallow, coarse, or overreach, adversely affecting related or general knowledge.
We introduceπͺππππππ β a general framework for Precise In-parameter Concept EraSure. π§΅ 1/
How can we interpret LLM features at scale? π€
Current pipelines use activating inputs, which is costly and ignores how features causally affect model outputs!
We propose efficient output-centric methods that better predict how steering a feature will affect model outputs.
New preprint led by my student @GurYoav with dream team @Roym4498, Chen Agassy, and Atticus Geiger π§΅1/