🚨📢Announcing the second Technical AI Governance Research (TAIGR) workshop @icmlconf. Accepting submissions (up to 8 pages) until April 24 on technical topics in AI governance! #icml2026
I'm extremely excited to be on the organizing committee this year for my favorite workshop ever!
Submissions (up to 8 pages) are due April 24! Co-submission with ICML and NeurIPS is encouraged!
https://t.co/40syiFt3Qk
🚨The 2025 AI Agent Index is out! 🚨
Amidst recent buzz over 🦀 and @NIST’s new agent initiative, we find:
- Selective reporting – esp. on safety
- Almost all agents backend just 3 model families
- Many agents don’t ID themselves as bots online
- Big US/China gaps
- And more…
🚨New paper🚨
From a technical perspective, safeguarding open-weight model safety is AI safety in hard mode. But there's still a lot of progress to be made. Our new paper covers 16 open problems.
🧵🧵🧵
New paper from the lab on how we need nuanced evaluation and discussion of LLM's cultural alignment. Their 'preferences' just aren't as stable, extrapolable, or steerable as past work suggests.
@aribak02@StephenLCasper@dhadfieldmenell
🚨New paper led by @aribak02
Lots of prior research has assumed that LLMs have stable preferences, align with coherent principles, or can be steered to represent specific worldviews. No ❌, no ❌, and definitely no ❌. We need to be careful not to anthropomorphize LLMs too much.
🚨 📣🧵 Introducing LAT-Chat -- https://t.co/zGMn0h33zc
We now have an interface to help red-team our Llama3-8b model trained using latent adversarial training. Happy jailbreak hunting! If you find a novel attack, consider sharing it via our form!
📢 Announcing the #SaTML2024 CNN Interpretability Competition results:
🥇 *New record* - Yun et al. - RFLA-Gen2
🥈 Moore et al. - FEUD
🥉 Nicolson - TextCAVs
👏 Tagade and Rumbelow - Prototype Gen
Report: https://t.co/FHJmSrd6x2
All were impressive!🧵 on new innovations below.
🚨 New paper: Defending Against Unforeseen Failure Modes with Latent Adversarial Training
We argue that LAT can be a key tool for safer AI because it can help address the gap between failure modes that developers identify 🎯 and ones they miss 🤔.
https://t.co/6xDUE0dWbX
🚨New paper🚨 8 Methods to Evaluate Robust Unlearning in LLMs
Unlearning is promising for safer LLMs, but it’s tricky to evaluate. Here, we (1) overview eval techniques, (2) red-team a popular method, and (3) show that ad-hoc evals can be misleading.
https://t.co/BX6RCzZbyN
🚨New paper🚨 Black-Box Access is Insufficient for Rigorous AI Audits
AI audits are increasingly seen as key for governing powerful AI systems. But to be effective, audits need to be high-quality, and to produce high-quality audits, auditors need access.🧵
https://t.co/SQw6aY4hdO
🧵📣Do LLMs "lie"? In a new EMNLP paper, we study what happens when what language models say seems to disagree with internal representations of truthfulness. The story is a little nuanced.
Thanks to coauthors Kevin Liu, @dhadfieldmenell, and @jacobandreas
https://t.co/cIl4Hq8fy2