π¨Are attention sinks a byproduct of optimization/training? Or are they sometimes functionally necessary in softmax Transformers?π¨
We prove that, in some settings, itβs the latter.
[https://t.co/iHfmDNWnxR]
π§΅
A sweet wrap-up to my visit at @Princeton: my work with @HazanPrinceton and Angelos Assos has been accepted to #NeurIPS2026! π
We show how to learn linear dynamical systems with a small memory footprint. Hereβs the idea π§΅
@Princeton@HazanPrinceton Experiments corroborate our findings: we construct a system where FIR, spectral filtering, and autoregressive predictors each fall short, while our combined predictor achieves far lower error with the same parameter budget.
A really neat idea IMO!
Intuitively, similar tricks could help Transformers acquire some of the abilities of recurrent and state-space models. Feels like a promising direction to explore!
New preprint: π₯ Latent Information Feedback Transformers (LIFT) π₯
Information in LM generation propagates downward (high to low layers) only through the decoded token, creating a bottleneck. Removing it via state propagation makes the model recurrent, which is not scalable for training.
Can we teach LMs to propagate state while keeping pretraining fully parallel? YES!
LIFT is a Transformer-based architecture + teacher-supervised training approach that enables Transformer LMs to exploit deep-to-shallow feedback, while keeping training parallel.
LIFT models consistently outperform standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under a token-matched budget, while being on par with or ahead of compute-matched Transformers. LIFT also outperforms Transformers trained on 8x more data on a state-tracking task, even when trained with the states of a Transformer that fails it!
Work by the amazing @TiroshDor and @AmosaurusRex
Paper: https://t.co/9AJFhJshkY
Detailed post + visualizations + code + models coming soon!
When and why does outcome-based RL yield step-by-step reasoning in Transformers?π¨
We investigate policy gradient on a graph-traversal task and formally characterize when and how reasoning emerges.
https://t.co/sODkMZNJDF
π§΅
When and why does outcome-based RL yield step-by-step reasoning in Transformers?π¨
We investigate policy gradient on a graph-traversal task and formally characterize when and how reasoning emerges.
https://t.co/sODkMZNJDF
π§΅
Robot policies can move but can't think.
LLMs can think but can't move.
So we connected them.
Real robot: 16.7% β 97.3%
Sim (LIBERO-PRO): 12.8% β 53.3%
Looped models scale compute by iterating in depth, demonstrating great potential in reasoning. But when should the looping stop, and how to keep it trainable at tens of thousands of unrolled layers? FPRM solves both. (1/7)
How can transformers memorize factual associations? It's common to think of MLPs as an associative memory, with parameters scaling linearly with # facts. We study an alternative: geometric factual recall. Joint work with @Giladude (eq. contribution), Joan Bruna and @albertobietti
Why AI agents fail to act safely on unseen tasks, even when task performance generalizes?π¨
Our recent work shows that generalizing safely is inherently hardβeven when agents succeed in trainingβmotivating research on new methods for agentic safety.
https://t.co/qTEt5xSzcH
π§΅
3 papers accepted at #ICML2026 π°π·π
πΈ From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
https://t.co/uSkeXBhmqe
πΈ Interpretability can be Actionable
https://t.co/SeoKTd3Bff
πΈHallucinations Undermine Trust; Metacognition is a Way Forward (out soon) @_galyo@ymatias
Congrats to all collaborators!
π§΅[1/3] Heading to #ICLR2026 π§π· to present our recent work on state space models:
The Curious Case of In-Training Compression of State Space Models, with the wonderful Makram Chahine, Daniela Rus and T. Konstantin Rusch.
ποΈ Fri, Apr 24 β’ 10:30 AM
π Pavilion 4 P4 - #5006
@NatalieShapira Great Q!
Luckly, we wrote another paper which aims to explains *why* this happens: https://t.co/yHUaSWnZli
Regurading *how*: we hope that this paper provides an answer for GPT-2, and hints that the answer might differ radically depending on architecture
π¨Are attention sinks a byproduct of optimization/training? Or are they sometimes functionally necessary in softmax Transformers?π¨
We prove that, in some settings, itβs the latter.
[https://t.co/iHfmDNWnxR]
π§΅
7/n
Huge credit to my coauthors Hila Ofek and Shahar Mendel, who were bachelor's students at the time and whom I had the privilege to mentor on this project. I'm very proud of you both, and very happy to see your hard work pay off!
Full paper at https://t.co/F2VWgoKeLx
π¨How are attention sinks formed in GPT-2?π¨
We uncover the circuit driving the first-token sink in GPT-2-style models, and use it to draw broader lessons about possible mitigations.
Excited to present this short paper at #ACL2026 .
https://t.co/F2VWgoKeLx
π§΅
6/n
This has direct implications for mitigation. Many pre-training changes may not generalize, and many post-training mitigations may need to be architecture-specific: targeting one component will help only if it is part of the active sink circuit in that model.