If you want to understand in less than 15 minutes why sparse attention emerges in Transformers, check out my NeurIPS 2025 oral!
https://t.co/V7HHNqX6sv
The first workshop on Computational Developmental Linguistics, collocated with ACL 2026, is welcoming submissions! Topics of interest include but are not limited to computational methods for developmental linguistics and language model learning dynamics.
Program details below:
At the #Neurips2025 mechanistic interpretability workshop I gave a brief talk about Venetian glassmaking, since I think we face a similar moment in AI research today.
Here is a blog post summarizing the talk:
https://t.co/LSwBf9XQzE
#NeurIPS2025 reflections
After a week of posters, hallway chats, and workshop deep dives, here are a few themes that stood out...
Training strategies and RL's evolving role: @YejinChoinka's keynote pushed past the "scale-is-all" narrative with a focus on RL--not just as a fine-tuning method, but as part of the pre-training story. The idea of eliciting reasoning behaviors in smaller models through RL is gaining traction.
Benchmark design and meta-evaluation: From Terminal-Bench to agent self-evaluation setups, we're seeing a shift toward evaluating not just outputs but internal reasoning and feedback loops. There's an emerging science of how we measure capability, not just completion.
LLMs in software environments: Papers like SWE-smith and SWE-rebench explored how LLMs perform in interactive dev environments. Robustness to tool changes and realistic regression testing stood out--especially relevant for agent-based use cases.
From scale to specificity: A subtle but important theme...quality, provenance, and realism in data pipelines (both human-labeled and synthetic) are being treated as first-class problems. That feels like a healthy turn from quantity-first mindsets.
Takeaway: NeurIPS isn't just about bigger models anymore--it's about sharper questions.
#NeurIPS #LLMs #AIresearch
An amazing blog post dropped on @huggingface explaining how today's LLM inference engines like @vllm_project work!
The concept of "continuous batching" is explained, along with KV-caching, attention masking, chunked prefill, and decoding.
Continuous batching is the idea of concatenating prompts and cleverly masking the attention to avoid wasting compute on padding tokens
1/2
Nested Learning might be the biggest AI shift in years.
It fixes the core flaw every model still has
They learn… and then forget.
Google’s new approach creates a layered memory system where some parts learn fast, others slow, so the model can update itself without destroying what it already knows.
If this scales, AI stops being a static snapshot and becomes something that actually grows with experience.
That’s the real step toward systems that learn continuously, not just predict.
New Anthropic research: Natural emergent misalignment from reward hacking in production RL.
“Reward hacking” is where models learn to cheat on tasks they’re given during training.
Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
📷📷📷New paper! (with @OpenAI) 📷📷📷
We trained weight-sparse models (transformers with almost all of their weights set to zero) on code: we found that their circuits become naturally interpretable! Our models seem to learn extremely simple, disentangled, internal mechanisms!
Can we find weight directions to modify LLM's behaviors?
Our new paper proposes contrastive weight steering, an alternative to activation steering for modifying behaviors using small narrow distribution data 🕹️
🧵👇
The Salon des Refusés posters are now online!
If you want your rejected NeurIPS paper featured, dont worry! You can still apply, but posters will be accepted on a rolling basis until the session is filled or Friday the 14th Nov, whichever comes first.
https://t.co/w8cu8EAIw0
📣NEW PAPER! What's In My Human Feedback? (WIMHF) 🔦
Human feedback can induce unexpected/harmful changes to LLMs, like overconfidence or sycophancy. How can we forecast these behaviors ahead of time?
Using SAEs, WIMHF automatically extracts these signals from preference data.
This paper draws connections between key concepts: attribution, unlearning, knowledge, FIM, and activation covariance, and proposes a very cool way to split up the weights of a model to separate memorization from generalization.
Insightful work @jack_merullo_!
Very excited to have been awarded a Google PhD fellowship in NLP for my work on mechanistic interpretability! Big thanks to @Googleorg for the support, as well as to my supervisors @sandropezzelle@boknilev and @ELLISforEurope for all the help along the way.
Can LLM verbalize a new concept? Yes (we plug it back in to evaluate)
Does LLM synonyms of such concept? Yes (and apparently *some* of them are common across models: machine synonym )
Can we use these new concepts to control and understand the model better? We think so, but you tell me. 🙂👇👇
https://t.co/If14uGX7jO
Based on neologism work - adding a word to mean a machine or human concept.
🧠 How can we equip LLMs with memory that allows them to continually learn new things?
In our new paper with @AIatMeta, we show how sparsely finetuning memory layers enables targeted updates for continual learning, w/ minimal interference with existing knowledge.
While full finetuning and LoRA see drastic drops in held-out task performance (📉-89% FT, -71% LoRA on fact learning tasks), memory layers learn the same amount with far less forgetting (-11%).
🧵:
I’m recruiting PhD students for 2026! If you are interested in robustness, training dynamics, interpretability for scientific understanding, or the science of LLM analysis you should apply. BU is building a huge LLM analysis/interp group and you’ll be joining at the ground floor.