1/9 🧭 Can an interpretable SAE feature also be used as a stable steering direction?
Our new preprint at @UKPLab provides geometric evidence that the answer is often no.
Project page: https://t.co/R0hfLpng7S
Thread 🧵
@UKPLab 8/9 Main takeaway:
An SAE feature can be interpretable and causally relevant without behaving like a stable steering direction.
Interpretability does not automatically imply steerability.
@UKPLab 6/9 Key result #2:
Value-like features (those encoding information such as factual attributes) more often produce structured, low-dimensional effects.
But those effects usually span several directions, not just one.
@UKPLab 5/9 Key result #1:
Consistent one-dimensional effects are rare across SAE variants.
Even when a feature is interpretable, its causal effect is often not captured by one reusable steering vector.
@UKPLab 4/9 We then analyze its geometry:
Does the feature act along one stable direction, a small set of directions, or a diffuse set of context-dependent directions?
@UKPLab 3/9 We introduce Feature-Effect Geometry Analysis, or FEGA.
Rather than asking only what information an SAE feature encodes, FEGA intervenes on that feature and asks: does it cause the same downstream effect across different contexts?
@UKPLab 2/9 Introducing:
“Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects” https://t.co/W1Zu1GF98e
We study whether SAE features produce consistent downstream effects across contexts.
We will introduce a novel dataset for the Vietnamese language (low-resource NLP community) at EACL2023! 🎊
ViHOS: Hate Speech Spans Detection for Vietnamese 🔍😠
📰https://t.co/Wox7u6SQaF
🔗https://t.co/JMcRjxgBvW
#NLProc#EACL2023@aclmeeting
With PhD applications due soon, thousands of young people are currently beginning their statements of purpose with the same cliché story, or the same anodyne statement
Stop right now!
Here are 10 thoughts for doing this right. Helps you, and helps admissions committees.👇
Zotero's inbuilt Note Editor can REVOLUTIONIZE your note-taking and writing processes.
But most academics don't know much about it.
Here's how to supercharge your writing using Zotero's Note Editor 👇
A step-by-step guide with visuals 🧵
One of the MOST CHALLENGING parts of any research project: Literature Review
Here's how to fast-track your literature review process using two FREE tools - Zotero and Research Rabbit 👇
A step-by-step guide with visuals 🧵
Every PhD student has to choose a supervisor.
But few know how to.
Joan Bolker, EdD, counseled PhD students for 30+ years at Harvard, MIT, and Brandeis.
Here is her advice on how to choose a supervisor:
Pretrained LMs are powerful! How to leverage their power for structured prediction? We are thrilled to share our paper “Autoregressive Structured Prediction with Language Models”, which achieves SOTA on NER, relation extraction, coreference resolution! #emnlp2022
15 years ago my PhD advisor taught me One Weird Trick for editing your own writing. Edit **back to front**, paragraph by paragraph. I still use it and it still surprises me how well it works. When I get my students to do it, it often blows their minds. Try it!
@suzan@TMLeiden@LIACS@UniLeiden@fhasibi Hi 👋,
I don't have Master degree, but I have a Bachelor one in Data Science and have some experience in researching NLP. Can I apply for this vacancy?