Latest research in Trustworthy ML. Organizers: @JaydeepBorkar @sbmisi @hima_lakkaraju @sarahookr Sarah Tan @chhaviyadav_ @_cagarwal @m_lemanczyk @HaohanWang
Our team in FAIR at Meta is hiring a (full-time) researcher!
We work on the topics of Reasoning, Alignment and Memory/architectures (RAM) for self-improvement & co-improvement.
Apply here:
https://t.co/Vukp3u8rfu
Location: NY, Seattle or Menlo Park.
Some of our recent work to give flavor:
Co-Improvement (position): https://t.co/XPwbsuCmSy
SPICE (Self-Play in Corpus Environments): https://t.co/47BarIqsFe
Self-Challenging Agents: https://t.co/qgDLmcgPjp
RL from Human Interaction: https://t.co/wmC2fVB0zu
AggLM (parallel aggregation): https://t.co/Fg0E31agT0
StepWiser (CoT-PRM RL): https://t.co/QbfBVYwxcu
DARLING (diversity-trained RL): https://t.co/J9ZSs8GnJp
J1 (RL-trained LLM-as-Judge): https://t.co/yG6xAPafTv
CoT-Self-Instruct: https://t.co/dHMYRxsXfJ
Multi-Token Attention: https://t.co/4kfUe8JQKl
The Terminal-Bench paper is here! Read it to learn where frontier models still fail and the secrets of how we sourced hundreds of high quality environments from our open source community. 🧵
Excited to share my work at Meta.
Knowledge Distillation has been gaining traction for LLM utility. We find that distilled models don't just improve performance, they also memorize significantly less training data than standard fine-tuning (reducing memorization by >50%). 🧵
This semester at UC Berkeley, I'm organizing an interdisciplinary reading group with @ProfHolliday and @bayesian_wasian on "AI Interpretability: Meaning, Methods, and Limits." Reading list: https://t.co/ACkiLOQgWi
Mechanistic interpretability occupies a unique role in AI safety discourse. Many people at frontier labs, including some who are deeply worried about the technology they're building, are investing significant hope in an ambitious research program: reverse-engineering their famously opaque AI models into human-understandable components.
With participants from Statistics, CS, Philosophy, Neuroscience, and more, we'll study state-of-the-art methods and ask what kinds of interpretability claims are actually meaningful, useful, and reliable.
We've structured the semester in three parts:
• Meaning: What is interpretability for? How do mechanistic explanations differ from post-hoc rationalizations?
• Methods: What can current techniques do? When does an intervention support an explanatory claim vs. merely provide a control knob?
• Limits: What are the fundamental barriers? What role can interpretability realistically play in reducing catastrophic risk?
We're still finalizing our reading list. If there's a paper that you think deserves more critical attention, or that changed how you think about interpretability, we'd love suggestions.
“I have a large, aligned, and safe model” — No One.
Had a great time speaking about Robust Unsupervised Probing Frameworks to evaluate the alignment of language models in the context of AI Safety.
Paper: https://t.co/Awbw4leRzY
Presenting this today at #ACL2025! Stop by if you’re interested in chatting about memorization and privacy! :)
Hall X5 Board #209 10:30-12
Hall X4 Board #259 16-17:30
🚨 Got a great idea for an AI + Security competition?
@satml_conf is now accepting proposals for its Competition Track! Showcase your challenge and engage the community.
👉 https://t.co/3g3nvv3yqa
🗓️ Deadline: Aug 6
I'm psyched for my 2 *different* talks on Friday @aclmeeting:
1.@llm_sec (11:00): What does it mean for an AI agent to preserve privacy?
2.@l2m2_workshop (16:00): Emergent Misalignment thru the Lens of Non-verbatim Memorization (& phonetic to visual attacks!)
Join us!
Super thrilled that HALoGEN, our study of LLM hallucinations and their potential origins in training data, received an Outstanding Paper Award at ACL!
Joint work w/i @shrusti_ghela*, and @davidjwadden@YejinChoinka 💫
L2M2 will be tomorrow at VIC, room 1.31-32! We hope you will join us for a day of invited talks, orals, and posters on LLM memorization. The full schedule and accepted papers are now on our website: https://t.co/AH3aMn3ID6
🚀 We introduce GrAInS, a gradient-based attribution method for inference-time steering (of both LLMs & VLMs).
✅ Works for both LLMs (+13.2% on TruthfulQA) & VLMs (+8.1% win rate on SPA-VL).
✅ Preserves core abilities (<1% drop on MMLU/MMMU).
LLMs & VLMs often fail because specific tokens (textual, visual, or both) push them toward hallucinations, bias, or unfaithfulness. Can we identify these tokens and steer model behavior without retraining?
⬇️🧵
I’m gonna be recruiting students thru both @LTIatCMU (NLP) and @CMU_EPP (Engineering and Public Policy) for fall 2026!
If you are interested in reasoning, memorization, AI for science & discovery and of course privacy, u can catch me at ACL!
Prospective students fill this form:
1/ 🚨New Paper 🚨
LLMs are trained to refuse harmful instructions, but internally, do they see harmfulness and refusal as the same?
⚔️We find causal evidence that 👈”LLMs encode harmfulness and refusal separately” 👉.
✂️LLMs may know a prompt is harmful internally yet still accept it!
We propose a ➡️harmfulness direction and a 🛡️Latent Guard model for #AI #Safety.
📢Happy to share that I'll join ELLIS Institute Tübingen (@ELLISInst_Tue) and the Max-Planck Institute for Intelligent Systems (@MPI_IS) as a Principal Investigator this Fall!
I am hiring for AI safety PhD and postdoc positions!
More information here: https://t.co/ZMCYXeC2fp
Had an amazing time @NewInML@icmlconf giving a talk on "What I Wish I knew before starting a PhD (but learnt the hard way)"!
Loved the post-talk discussions and the heart warming messages :)
Sharing slides since some people asked, link in the tweet below 👇
🧵 Academic job market season is almost here! There's so much rarely discussed—nutrition, mental and physical health, uncertainty, and more. I'm sharing my statements, essential blogs, and personal lessons here, with more to come in the upcoming weeks! ⬇️ (1/N)
We are seeking emergency reviewers for the NeurIPS 2025 ethics review. If you are interested and available to contribute this week, please sign up at https://t.co/UGlrqcOHwc.
@NeurIPSConf@trustworthy_ml@acidflask#AI#ML#NeurIPS