Excited to share that our paper, “Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics”,
has been accepted to NeurIPS 2026!
Work done with Matt Smart, Soumya Ganguly, Nilava Metya and Alexandre Morozov.
arXiv: https://t.co/R9axqnsvCC
I'm looking forward to giving a talk tomorrow morning at the ICML workshop on High-Dimensional Learning Dynamics (HiDL) https://t.co/YjnAgcjvyw. Come by at 9 am!
Sad to miss #ICML2025 this time; go check out the awesome @rdMorel presenting our DISCO model. It learns the evolution operator from short trajectory data without knowing physics and stays SOTA on next-step prediction! Poster W-107, this Wed (July 16)!
Great to see this one finally out in PNAS! Asymptotic theory of in-context learning by linear attention https://t.co/LFube48Bnl Many thanks to my amazing co-authors Yue Lu, @maryiletey, Jacob Zavatone-Veth and @AninditaMaiti7
We also show how this attention model relates to a one-step update of a modern Hopfield network, clarifying a potential connection indicated by Ramsauer et al. (https://t.co/Eevc4UpZZi…)
At #ICML2025, presenting work done with Matthew Smart and @albertobietti on in-context denoising (https://t.co/4jj5CRTtDS). Matthew's oral is on Thursday, 4:15-4:30 PM, at West Ballroom A, and our poster #E-3207 is presented on Thursday, 4:30-7:00 PM, at East Exhibition Hall A-B.
We investigate how a one-layer transformer can solve an in-context denoising problem by averaging over samples from a distribution with the right attention weights.
We also show how this attention model relates to a one-step update of a modern Hopfield network, clarifying a potential connection indicated by Ramsauer et al. (https://t.co/Eevc4UpZZi…)
ICML this week! Come by
T PM @LauditiClarissa's work on muP BNNs https://t.co/lkpxI3rqlT
W AM, model of place field adaptation@mgkumar138, Jacob ZV https://t.co/FB6AQ5fRfF
W PM a model of LR transfer in linear NNs https://t.co/JINZhapMz5
all from senior author @CPehlevan!
Nice article! I appreciate that it mentions my work and the work of my students. I want to add to it.
It is true that there is some inspiration from spin glasses, but Hopfield is much bigger than spin glasses. The key ideas that resurrected artificial neural networks in 1982 were: collective and distributed computation. Individual compute units (neurons) comply with only local rules. Despite that computation EMERGES at the network level. We can kill individual neurons and connections (up to an extent) and the network will continue performing computation. The ideas of emergence (things that you don’t put into the system “by hands” but that appear spontaneously by means of interactions) are ubiquitous in Physics - far beyond a relatively small field of spin glasses. Physicists LOVE emergence and have some of the best tools to study it.
Some of the ideas that were credited by the Nobel Prize existed before. What was remarkable about Hopfield and Hinton is that they had good taste, right vision, and relentlessly executed on their vision despite community at large believing that they were wrong.
I find it amusing when I speak with some of my colleagues from more traditional physics departments and hear comments like “hmmm, this is nice, but is this really Physics?”. The field of Physics of Computation was created by R.Feynman, J.Hopfield, and C.Mead in the 1980s. Back then it was not anything controversial that THIS IS PHYSICS. Fast forward to 2025, AI is running a trillion dollar economy and impacting every aspect of our lives, some of the most consequential experiments in human history are being run in data centers around the world, the Nobel Prize in Physics has been awarded to artificial neural networks. How come we still wonder if this is Physics?
My humble prediction: those departments and companies that embrace Physics of Neural Computation and invest in it will thrive, and those who don’t will become obsolete very soon.
Come hear Matt Smart's talk about in-context denoising with transformers at the Associative memory workshop #ICLR25, 2:15pm! This task refines the connection between transformers and associative memories. w/ M Smart and @AnirvanMS at @FlatironInst
Paper: https://t.co/O1joW5BUd8
@SuryaGanguli Done! Also, I don’t think the current madness is just a result of left’s overreach (however irritating). The trouble is that part of the madness, especially the attack on academia, there is method in’t.
There are days in life that shake you.
I’m shattered 💔 to share that I just found out that the US Government terminated my 2024 NIH Director’s Early Independence Award (~$2 million), threatening my long-promised assistant professor job at @Columbia & academic career... 1/🧵
Learning important spatiotemporal features from biological images. Had fun working with Siddhartha Saha, Qixin Yang, @WLosert and @AlexandreMoroz3
https://t.co/RaH3YfOvNf
Excited to share our work on ‘Deep Learning Based Superconductivity Prediction and Experimental Tests’ at the #NeurIPS2024 workshop on Machine Learning and Physical Sciences.
https://t.co/05CyyZGz6A
Want to do ML research in NYC with lots of freedom and no drama? Apply here to be a @FlatironCCM research fellow (deadline Dec 15): https://t.co/na5IlPwRzh