Can we give Transformers a rich deep-to-shallow feedback channel across steps, without sacrificing parallel pretraining?
This was the main focus of my MSc in @megamor2’s lab, and I’m happy to share what we found.
More results, code, and models coming soon. Stay tuned! ⏰
New preprint: 🔥 Latent Information Feedback Transformers (LIFT) 🔥
Information in LM generation propagates downward (high to low layers) only through the decoded token, creating a bottleneck. Removing it via state propagation makes the model recurrent, which is not scalable for training.
Can we teach LMs to propagate state while keeping pretraining fully parallel? YES!
LIFT is a Transformer-based architecture + teacher-supervised training approach that enables Transformer LMs to exploit deep-to-shallow feedback, while keeping training parallel.
LIFT models consistently outperform standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under a token-matched budget, while being on par with or ahead of compute-matched Transformers. LIFT also outperforms Transformers trained on 8x more data on a state-tracking task, even when trained with the states of a Transformer that fails it!
Work by the amazing @TiroshDor and @AmosaurusRex
Paper: https://t.co/9AJFhJshkY
Detailed post + visualizations + code + models coming soon!
Information flow in auto regressive LLMs is mostly bottom-up - except for a single decoded token.
Allowing information from deep layers to propagate to the next step as a latent state is attractive, but hard to train due to the implied recurrence.
In our recent preprint @TiroshDor does so efficiently! by finding that existing LLMs can be used to teacher-force the latent state, allowing parallel training and a stationary distribution for the states.
The result is LIFT, a Transformer LLM that carries a states from deep-to-shallow layers. LIFT outperform the baselines in all settings considered under a token-matched budget and is on-par or ahead under matched compute.
Happy to be part of this project with the great @megamor2 and @TiroshDor !
Paper: https://t.co/nTejQ9Ncq9
A really neat idea IMO!
Intuitively, similar tricks could help Transformers acquire some of the abilities of recurrent and state-space models. Feels like a promising direction to explore!
Three papers from our lab at #NeurIPS2026, spanning LLM reasoning faithfulness, belief formation, and interpretability of MoEs:
⭐ Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
https://t.co/IxmJMz4Uyb
We introduce a benchmark for CoT faithfulness metrics, and find that existing metrics are either close to random or prohibitively expensive to run
@GurYoav@anmarasovic
⭐ Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
https://t.co/S6hl0PrtgI
An interdisciplinary collaboration applying the HOT-3 indicator of consciousness to LLMs using interpretability tools, revealing surprising patterns in belief formation and action selection in LLMs
@noam_steinmetz@GoldsteinYAriel@Liad_Mudrik
⭐ Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
https://t.co/RepFOlWAne
We find a geometric coupling between routers and their corresponding experts in MoEs, explaining how routers form assignments that support an effective division of labor
@sageaarc@noyahochwald
We didn’t explore combining LIFT with speculative decoding in this paper. My intuition is that a draft model could propose both tokens and the states passed between steps, instead of tokens alone. Exploring the impact on draft acceptance rates and overall decoding speedup would be an interesting direction for future work.
@Zhen4good@megamor2 Exactly! On context length: we see gains across our small, medium, and large models, using contexts of 1,024, 2,048, and 4,096 tokens, respectively. It would be interesting to test whether these gains persist at even longer contexts.
@cranialxix@megamor2 Thanks for sharing Maglev!
The jointly trained prefiller and parameter sharing are particularly interesting. It would be interesting to explore how these choices interact with LIFT’s distribution feedback.
Thanks! I see the approaches as potentially complementary: NextLat encourages more predictive representations through a next-latent prediction objective, while LIFT changes information flow by feeding deep representations back into the LM’s input at the next step.
Combining the two would be an exciting experiment we might explore. My intuition is that richer representations could be even more useful when lower layers can directly access them at subsequent steps.
@megamor2@noam_steinmetz Really happy to finally share this work! 🔥
Huge thanks to @megamor2 and @AmosaurusRex 👑 👑, it was amazing working on this together.
Code, models, and more details coming soon! ⏰
🗣️🧠 Speech Language Models require lots of compute to train, right?
In our new paper, we test is it possible to train an SLM on 1xA5000 gpu in 24 hours?
The results may surprise you (they even surprised us)!
Tips, open source resources, full paper 👇🏻