New preprint, w/ @neuro_kim! We characterize the compositional generalization behavior of kernel models. Our theory derives a new compositional generalization class, highlights key failure modes, and is empirically valid for deep neural networks. (1/24)
https://t.co/1oqBLgKR2R
Lots more details on this are in the paper! We’re very excited about this direction and are actively working on follow-ups. If you’re curious about any of this, please come check out the poster on Friday (or reach out to us in any other way :) ). (13/13)
https://t.co/VO06HtrJlU
I’ll be presenting my joint work with @Jack_W_Lindsey on the inductive biases of finetuning and multi-task learning at @NeurIPSConf this Friday. Since the paper has changed substantially over the last year, it’s time for an updated tweeprint! (1/13)
https://t.co/DUMHqQESTD
When do transformers length-generalize?
Generalizing to sequences longer than seen during training is a key challenge for transformers. Some tasks see success, others fail — but *why*? We introduce a theoretical framework to understand and predict length generalization.
How does the brain bind action to value?
How does it navigate a tradeoff between this binding and generalization to novel situations?
Out today in Nature Neuro! https://t.co/YPiY2uHoFz
@JustFineNeuro @NeuroPolarbear @BecketEbs@neuromochi
See original 🧵 + updates below
We live with the limitations of our memory, but don’t really know where they come from. Our new paper (https://t.co/DYiL0IfxvF) studies "swap errors", which we argue arise during memory manipulation – see thread for more! @timbuschman, @MattPanichello, @wjeffjohnston
How do learning systems implement relational and compositional generalizations? In our new preprint (w/ Kenny Kay, Greg Jensen, @vferrera, and Larry Abbott), we investigated this question using transitive inference as a case study. (1/13)
https://t.co/hosyparaYF
Thanks to the reviewers who gave us a lot of helpful feedback and in particular encouraged us to dive deeper into why feature-learning neural networks fail to generalize on transitive inference!
Excited to announce (a bit belatedly) that this is now out @PNASNews. If you're interested in how standard statistical learning principles can enable relational generalization, check out our article here: https://t.co/OchSIEQUs9
How do learning systems implement relational and compositional generalizations? In our new preprint (w/ Kenny Kay, Greg Jensen, @vferrera, and Larry Abbott), we investigated this question using transitive inference as a case study. (1/13)
https://t.co/hosyparaYF
Lots more details on this are in the preprint. We’d love to hear any thoughts on this work and are excited about exploring these phenomena further in future work! (24/24)
https://t.co/1oqBLgLoSp
Taken together, conjunction-wise additivity may be a useful compositional generalization class: on additive tasks, analyzing kernel models can help us understand behavior in deep networks; on non-additive tasks we need to investigate other learning mechanisms. (23/24)