$1,000,000 to understand how LLMs write code.
Announcing: The Martian Interpretability Challenge.
Understanding the inner workings of LLMs is the greatest scientific challenge of our age,. Let's solve it.
Apply here: https://t.co/taai9PSu6d
🧵👇
have lots of unlabeled data and want to use it to improve evals?
come talk to me and @shuvom_s today at 4:30PM, Poster #106 at NeurIPS! I'll also be around through Saturday, do reach out to chat!
I had a ton of fun and learned a lot working on this with my many amazing collaborators from MIT and MGH, including senior author @MazAbulnaga!
Paper link: https://t.co/ASU2DNlqwi
Excited to be in San Diego for #ML4H2025 and #NeurIPS2025 this week! DM me to connect!
I’m presenting a poster at ML4H this afternoon. See 🧵for details and to learn what’s in the teaser image (hint: it’s not feminist art)
Our method works by combining ordinal regression (to account for ordinal labels) with soft label training (to account for annotator disagreement)
While here we focus on phonotrauma, our approach is applicable to any problem with ordinal targets and multi-rater annotations
I am on the job market this year! My research advances methods for reliable machine learning from real-world data, with a focus on healthcare. Happy to chat if this is of interest to you or your department/team.
Very excited to be in Phuket, Thailand for #AISTATS2025 🇹🇭
DMs are open if you want to connect and chat!
I’ll be presenting my recent work with amazing collaborators @ShenRaphael and @ta_broderick on extending multimarginal Schrödinger bridges using domain knowledge 🧬🦠🌀🌊
Check out the paper to learn more! https://t.co/NwIC5BBvIW
A huge thanks to my awesome collaborators @osazuwa and @emrek at @MSFTResearch and John Guttag at MIT - I learned so much from working w/ you on this!
We conducted experiments on:
- 2 QA tasks (social bias and medical QA 🩺)
- 3 LLMs (Claude-3.5-Sonnet, GPT-3.5, GPT-4o)
This analysis revealed new insights into the ways in which LLMs are unfaithful, such as....
On a medical QA task, LLM explanations contain misleading claims about which pieces of evidence are influential.
e.g., on this question, "the patient's mental status" has the largest effect on Claude's answers, yet the LLM *never* mentions it in its explanations
The key challenge is estimating the causal effects of concepts. To do this we:
(1) Use an auxiliary LLM to generate realistic counterfactuals w/ edited concept values
(2) Use a Bayesian hierarchical model to capture the effects of concepts at both the dataset and question level
Excited to present our ⭐️spotlight⭐️ paper on measuring the faithfulness of LLM explanations at #ICLR2025! I'll be at the 3pm poster session today 4/25: https://t.co/Hren32QlMi. Details in 🧵
Also, I'd love to chat interpretability, causality, or representation learning (DM me)!
To address this, we introduce causal concept faithfulness, which compares:
🗨️ THE TALK: which high-level concepts in the question (e.g., gender) does the LLM say affected its answer?
🚶♀️THE WALK: which concepts actually have a causal effect on the LLM's answers?
To make LLMs safer for users, we'd like to be able to identify:
1⃣ the degree of unfaithfulness 📏
2⃣the specific *ways in which* the LLM's explanations are misleading (e.g., masking gender bias) 🔎
Prior faithfulness metrics are designed for 1⃣ but they don't accomplish 2⃣
LLMs can give plausible explanations of their answers to questions🤖
But their explanations can misrepresent the LLM's true decision-process, aka they can be unfaithful!🤥
e.g., an LLM says it prefers a job candidate due to their skills🖊️ but it actually chose based on gender⚧️