@PalantirTech Rifles are the wrong metaphor. The closest historical metaphor for the level of personalization and scope of AI militarization is secret police.
[Hot take] Your Causal Variables Are Irreducibly Subjective
Mech interp keeps "finding the bug" in earlier interventions, but the real problem is upstream: your variable definitions are subjective choices no formalism can validate.
https://t.co/PUkIlUQgIv
Democracy depends on an informed electorate. But political issues and ballot measures can be confusing, obscuring the effects of one outcome versus another. Moreover, politics is personal. Once we make an initial decision about an issue, it can be hard to change our mind or see things from “the other side.” And talking about issues with those with whom we disagree can be challenging, especially when the conversation feels more like a debate than a discussion.
Technology offers ways to alleviate these difficulties, but not without introducing problems of its own. The Internet and social media promised new ways for people to connect, discuss issues, and learn from each other. But in practice, both often inflame passions, solidify echo chambers, and spread misinformation. More recently, LLM chat interfaces may help people stay informed through personalized access to information, but mainstream chatbots tend to match user beliefs rather than clarifying or challenging them.12 Without the kind of pushback you’d encounter in a discussion between disagreeing friends, chatbots are ill-suited for helping people think through political issues in a balanced way.
The goal of CivicChats is to address these shortcomings. Starting with ballot measures, CivicChats helps people better understand political issues through three different modes of discussion: a Q&A mode for understanding what a measure does and what’s at stake, an argumentative mode that presents competing views to your own, and a reflective mode that helps you examine and develop your own thinking.
Most mech interp work relies on activation patching, but patching activations destroys previous computation.
What if we want to use a different mechanism on the same residual stream? We propose dynamic weight grafting to interpret finetuned model weights.
🧵 1/n
Predictive Interpretability > Mechanistic Interpretability
Prompting is the best method of scientific inquiry we have to study LLMs
It's socially devalued because it doesn't include much d/dx,O(),etc.
come to poster #3503 to talk about this or anything re: the science of LLMs
Excited to present our work on LLM-assisted explainability at #ICML2025!
🖼️ Poster: Wednesday, 11:00am–1:30pm (#E-2902)
📄 https://t.co/8YVJksPjqi
w/ @seanrson @toddknife@ggarbacea@victorveitch
If you're using LLMs to generate counterfactual pairs, rewrite twice—not once!
Semantics in language is naturally hierarchical, but attempts to interpret LLMs often ignore this.
Turns out: baking semantic hierarchy into sparse autoencoders can give big jumps in interpretability and efficiency.
Thread + bonus musings on the value of SAEs:
1/n You may know that large language models (LLMs) can be biased in their decision-making, but ever wondered how those biases are encoded internally and whether we can surgically remove them?
5/ Shout out to my amazing collaborators @seanrson @toddknife@ggarbacea@victorveitch give them a follow!
Dive in to the full paper at https://t.co/8YVJksPRfQ
🧵 RATE: Score Reward Models with Imperfect Rewrites of Rewrites
1/ How do you measure whether a reward model incentivizes helpfulness without accidentally measuring length, complexity, etc?
Rewrites of rewrites give good counterfactuals, without needing to list all confounders!
4/ But how can we know we’re actually getting counterfactuals? We use a novel synthetic experiment to test how much our rewrite method is affecting known off-target correlates: induce a distributional shift, and see if the reported ATE changes! (It shouldn’t, and RATE passes ✅)
Reward models are a key ingredient to understanding LLMs. But what they *actually* reward can be a mystery.
Turns out we can measure this directly! The trick is to use LLMs to produce "rewrites-of-rewrites" datasets to measure attributes in isolation.
https://t.co/yjtog1XB4h
Are LLMs just doing next token predictions? It is believed that if an LLM can accurately predict the next tokens in a Wikipedia entry, it essentially "learns" the information.
But do pre-trained LLMs actually need to understand context sentences to solve this task? The answer is no!
Progress in empirical ML fields like interpretability is driven by success metrics -- but how good are these metrics? @JosephMiller_ et al find "faithfulness" measures are sensitive to arbitrary design choices, calling into question previous interp claims.
Fundamentally, high-level concepts group into categorical variables---mammal, reptile, fish, bird---with a semantic hierarchy---poodle is a dog is a mammal is an animal.
How do LLMs internally represent this structure?
https://t.co/HK2iFLUpte
LLM best-of-n sampling works great in practice---but why?
Turns out: it's the best possible policy for maximizing win rate over the base model!
Then: we use this to get a truly sweet alignment scheme: easy tweaks, huge gains
w @ybnbxb@ggarbacea
https://t.co/GHaP911Lil
Ongoing work, joint with @ggarbacea and @victorveitch.
-> We formalize mechanistic faithfulness as structural equivalence between causal abstractions
-> *Partial* faithfulness interp evaluations are naturally induced by causal hierarchy
'Mechanistic' interpretability has unclear objectives, inconsistent evaluation, and overlooks alternative hypotheses.
Come discuss a causal taxonomy for interp evals! CRL workshop poster session: Friday Dec 15, 10:30-12pm in 243-245! @crl_neurips2023
https://t.co/grUFLjbVm0