How do you know contrastive probing data identifies a unique feature?
How can we identify directions that model combinations of features?
We propose Contrastive Eigenproblems to tackle both of these issues.
Come see the poster at the MechInterp Workshop @ NeurIPS this sunday!
In the paper, we also explain: how Contrastive Eigenproblems take inspiration from, and are explanatory of, Contrast-Consistent Search, a well-known probing method, and (2) why COPA does not isolate one direction. Plus some more theoretical results.
https://t.co/gtErbg7W3W
How do you know contrastive probing data identifies a unique feature?
How can we identify directions that model combinations of features?
We propose Contrastive Eigenproblems to tackle both of these issues.
Come see the poster at the MechInterp Workshop @ NeurIPS this sunday!
We can also use more types of contrast to solve a single Contrastive Eigenproblem. When applied to the 'sentence polarity' and 'sentence truth' features, the top eigenvectors yield directions for both, plus the polarity-sensitive truth direction first found by Bürger et. al.
What can we learn about how named entities are represented in language models by substituting them for different ones and observing how predictions change?
Come see my poster this Thursday December 8th @ BlackboxNLP 2022 in Abu Dhabi (and online).
@norabelrose Yes interesting idea. You'd def have to do it for a bunch of local optima to make sure it's not a fluke. Also seems related to counterfactual explanations, except instead of a cfactual input you try to find cfactual weights that still perform well on main task but not on aux.
This leads us to question the underlying assumption that we should measure the validity of attention-based explanations based on how well they correlate with existing feature attribution explanation methods.
Very excited that our paper 'A Song of (Dis)agreement: Evaluating the Evaluation of Explainable Artificial Intelligence in Natural Language Processing' w/ Michael Neely, @MauritsBleeker and @__alucic got accepted for @hhai_conference#HHAI2022
https://t.co/yd4g1oPB8i
We find that attention-based explanations do not correlate strongly with any recent feature attribution methods, regardless of the model or task.
Furthermore, we find that none of the tested explanations correlate strongly with one another for the transformer-based model.
1/14 Whether Russians actually support the hideous war that Putin has waged against Ukraine is a matter of utmost political importance. The answer to this question will largely define Russia’s place in the history of the 21st century.