1/9: Dense SAE Latents Are Features💡, Not Bugs🐛❌! In our new paper, we examine dense (ie. very frequently occuring) SAE latents. We find that dense latents are structured and meaningful, representing truly dense model signals.🧵
Meet Ai2 Paper Finder, an LLM-powered literature search system.
Searching for relevant work is a multi-step process that requires iteration. Paper Finder mimics this workflow — and helps researchers find more papers than ever 🔍
What can AI researchers do *today* that AI developers will find useful for ensuring the safety of future advanced AI systems? To ring in the new year, the Anthropic Alignment Science team is sharing some thoughts on research directions we think are important.
LLMs have behaviors, beliefs, and reasoning hidden in their activations. What if we could decode them into natural language?
We introduce LatentQA: a new way to interact with the inner workings of AI systems. 🧵
NeurIPS has an overwhelming amount of papers, so I made myself a hacky spreadsheet of all (well, most) of the interpretability papers - sharing in case others find it useful!
It's definitely got false negatives and positives, but hopefully is better than baseline.
This paper on the statistics of evals is great (and seems to be flying under the radar): https://t.co/80cxSYbvQk
The author basically shows all the relevant statistical tools needed for evals, e.g. how to do compute the right error bars, how to compare model performance, and how to do power analysis.
Back when @jeremy_scheurer and I wrote the "We need a Science of Evals" post (https://t.co/zs1rJHtdr7) this paper is exactly the kind of thing we had in mind and more.
Our paper on individual neurons that regulate an LLM's confidence was accepted to NeurIPS! Great work by @alesstolfo and @benwu_ml
Check it out if you want to learn about wild mechanisms, that exploit LayerNorm's non-linearity and the null space of the unembedding, productively!
New paper w/ @benwu_ml and @NeelNanda5!
LLMs don’t just output the next token, they also output confidence. How is this computed?
We find two key neuron families: entropy neurons exploit final LN scale to change entropy, and token freq neurons boost logits proportional to freq 🧵
📢 We're happy to announce that our group has 12 papers accepted to #EMNLP2023 (6 main, 6 findings) 🧵⬇️
Congratulations to all our members and collaborators! 🥳 #NLProc