Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures?
Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham
https://t.co/vIdDx3FdL9
New preprint from the lab!
Behavioral and neural mechanisms for stochastic choices in mixed-strategy games: https://t.co/48OlCsCwiW
We use game theoretic tasks, cross species comparisons, modelling strategies, and mesoscale imaging to probe how we can escape predictability!
Version 2 of Theory of Contravariance is out! https://t.co/HEQpw3qprl New material on contravariance for Transformers, and the theory of Representational Similarity Analysis (RSA) and centered kernal analysis (CKA).
For transformers: it turns out that they have privileged axes, just like convnets, if you look in the right place (MLP layers and attention heads.) The identification of privileged heads is a potentially key result for emergence of interpretable stucture in LLMs. @meenakshik93
For RSA: it turns out that you can decompose RSMs into unique task-relevant "core geometry" and a task-irrelevant symmetry-generated term. Weak-strong equivalence holds for the core geometry and by projecting onto privileged axes you can filter out the task-irrelevant part so it doesn't interfere. This builds on work from Marvin Theiss @sciencelukas@saxelab@ermgrant
The Principle Investigator will do a deep dive on both topics in the coming days! https://t.co/xCFbmgSz48
Presenting your work with friends is always nicer!
Make sure to checkout Neuralplayground published as part of @CogCompNeuro proceedings!
https://t.co/HkCDJjqla8
The amount of basic research done in industry is nowhere near the amount done in universities.
At various times, there have been industry labs that had fundamental research activities and have made important scientific contributions.
Examples in information technology include Bell Labs, IBM Research, Xerox PARC, GE, Phillips, NEC, and several other. That disappeared in the 1990s.
Microsoft Research picked up the torch in the 2000s, followed (to some extent) by Google and then Meta (for about a decade until recently).
But their innovations almost always built on top of academic work, and certainly profited from the whole research ecosystem.
This is the very first project I started working on during my PhD with the amazing PPSleep team, and I’m so excited to finally see it published! 🎉
🧠Replay of procedural memory is independent of the hippocampus
https://t.co/EXEVJiOktf
New paper w/ @SaxeLab & Nishil Patel! 🧵
Does RL post-training teach models anything new, or just amplify skills already in the base model?
We built a fully auditable testbed to settle it — and caught RL composing new strategies in the act.
A question on synthetic data generation: If we want a language model to solve k-step arithmetic problems (such as a+b*c-d=?), with operands from 1 to 100, which training distribution should we use?
A. Uniform distribution: Sample these k operands uniformly from 1 to 100
B. Power law: randomly shuffle 1-100 and impose an artificial power law. Sample these k operands according to this power law.
⚡Our ICML 2026 (spotlight) paper shows: Option B is better! Surprisingly, the same idea extends far beyond this simple example to many reasoning tasks that require implicit composition of multiple atomic skills, including multi-hop QAs and synthetic GSM problems.
📄Paper: https://t.co/8LVlmCCBz9
📝Blog: https://t.co/t2yhIJ0s1P
Pretraining + fine-tuning powers modern ML, but we lack a theoretical understanding of how pretraining actually shapes downstream learning.
Enter our new @icmlconf paper: “A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning”
📅 July 9th, Poster #4502 Session 8!
🧵
Interested in any of our works? Come and say hi @icmlconf!
Unfortunately I won't be around this year but my collaborators will answer all your questions :)
Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization.
Europe has a lot to lose in the current AI race, and it's worth examining how threats to middle-power sovereignty can result in unsafe outcomes.
Such scenarios help illustrate why Europe must invest in AI initiatives that can either leapfrog the current frontier or offer critical components like safety and reliability.
Model collapse is often framed as “models getting worse”
In our ICML Spotlight Position paper, we show a high risk of unequal degradation. Rare languages, minority viewpoints, and low-resource communities are likely to be affected first and most severely
https://t.co/Sm1aESb85m
I’m excited to share that our paper “Compositionality and systematicity emerge from iterated learning in deep linear networks” has been published at PNAS. This work was conducted with @kleinric@BenjaminRosman and @SaxeLab. Some highlights below.
https://t.co/yKJRpUT0Qk
1/ Deep learning is going to have a scientific theory. We can see the pieces starting to come together, and it's looking a lot like physics!
We're releasing a paper pulling together these emerging threads and giving them a name: learning mechanics.
🔨 https://t.co/92nSIHameW 🔧
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures?
Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham
https://t.co/vIdDx3FdL9
Two Analytical Connectionism-related updates:
1. ⏰ 1 week left to apply! Interested in language + AI & cognition? Don’t miss it: https://t.co/pBuHTe2Hm3
2. 📜 Lecture notes from the first two editions are finally out: https://t.co/fzc8E7Yx3G
Postdoc opening!
Come work with us on deep learning theory relevant to AI safety
Deadline: 7 Apr 2026
Details and application: https://t.co/WeyecoXWhk+
Very excited by this year's Analytical Connectionism Summer School!
A dream lineup of speakers on the topic of language acquisition in minds and machines
Bursaries available to cover costs
Aug 17 – Aug 28, 2026 Gothenburg
Details: https://t.co/fP4HYwRF9C
Looking for alternatives to quadratic functions for closed-form analysis in optimization? This post explores matrix Riccati dynamics and their applications to neural networks. https://t.co/nEJdNGnNEW