4/4
Our results suggest safety should be modeled on its own, and not assumed to emerge as a byproduct of performance generalization.
Read more in our paper!👇
https://t.co/qTEt5xSzcH
Joint work w/ Yotam Alexander, @tomerslor, @yoav_nagel and @nadavcohen
Why AI agents fail to act safely on unseen tasks, even when task performance generalizes?🚨
Our recent work shows that generalizing safely is inherently hard—even when agents succeed in training—motivating research on new methods for agentic safety.
https://t.co/qTEt5xSzcH
🧵
3/4
We empirically demonstrate this hardness in linear-quadratic control, quadcopter navigation and LLM-based agentic CRM.
Notably, while safe and unsafe agents exhibit comparable performance on tasks seen in training, a large gap emerges on unseen tasks.
🚨Are attention sinks a byproduct of optimization/training? Or are they sometimes functionally necessary in softmax Transformers?🚨
We prove that, in some settings, it’s the latter.
[https://t.co/iHfmDNWnxR]
🧵
I’m at @IMSI_Institute this spring as a visiting researcher for the Long program on Theoretical Advances in Reinforcement Learning and Control. Excited to meet new people and explore the intersection of deep learning theory and RL/control.
When and why does outcome-based RL yield step-by-step reasoning in Transformers?🚨
We investigate policy gradient on a graph-traversal task and formally characterize when and how reasoning emerges.
https://t.co/sODkMZNJDF
🧵
📜🚨
Introducing TensorLens! 🔎
Our new tool for Transformer & LLM interpretability.
The problem: attention matrices are (i) a shallow view that ignores embeddings, FFNs, and values, and (ii) there are too numerous (per head & layer), which quickly becomes overwhelming.
🧵 1/6
At #NeurIPS? Check out posters by my fantastic students Yotam Alexander, @YoniSlutzky, @YuvalRanMilo :
Wed 4:30–7:30pm, hall CDE: The Implicit Bias of SSMs Can Be Poisoned With Clean Labels
Fri 11:00am–2:00pm, hall CDE: Do NNs Need GD to Generalize? A Theoretical Study
Do neural nets really need gradient descent to generalize?🚨
We dive into matrix factorization and find a sharp split: wide nets rely on GD, while deep nets can thrive with any low-training-error weights!
https://t.co/rfsepXiEQP
🧵
How can the implicit bias of SSMs fail, even with clean labels? 🚨
Like many deep learning models, SSMs trained with GD are often implicitly biased to fit data in ways that generalize well. But we show this bias can be poisoned by cleanly labeled data!
https://t.co/mvzM5DwLYs
🧵
5/5
Overall, our findings suggest that even in the simplest setups, there’s no straightforward answer to whether neural nets really need GD to generalize well.
Read more in our paper!👇
https://t.co/rfsepXiEQP
Joint work w/ Yotam Alexander, @YuvalRanMilo and @nadavcohen
Do neural nets really need gradient descent to generalize?🚨
We dive into matrix factorization and find a sharp split: wide nets rely on GD, while deep nets can thrive with any low-training-error weights!
https://t.co/rfsepXiEQP
🧵