Introducing PC-ALM, a local-learning alternative to backpropagation.
Our method trains 1000-layer neural nets using only local dynamics, and without backprop.
Blog: https://t.co/bBGCgalqKW
Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly. How can a physical system, such as the brain, solve multilayer credit assignment without explicit use of backprop?
We look for inspiration in two related fields: distributed optimization and NeuroAI.
In NeuroAI, predictive coding asks each neuron activation to solve an energy-based inference problem instead of using a standard forward pass. That inference step can be implemented as energy-minimization dynamics on local prediction errors.
This perspective -- each layer as a dynamical system -- has proven promising, but performance of predictive coding hasn't scaled well with depth. Credit signals at far ends of the network struggle to diffuse into internal layers.
We turn to distributed optimization, generalizing predictive coding to use an augmented Lagrangian instead of energy. This motivation stems back to a classic 1988 paper by LeCun, showing that the Lagrange multipliers of a deep network can be identified with gradients of a supervised loss. The augmented Lagrangian then bridges LeCun's perspective to the standard predictive coding that is used in NeuroAI.
We find that this new perspective yields a natural PC-like alternative to backpropagation, resulting in a method we call PC-ALM. PC-ALM differs from PC in that it introduces dual neurons (Lagrange multipliers) as part of the layer-local dynamics, resulting in each layer acting as a PI feedback control system to minimize local prediction errors.
We find that PC-ALM is capable of propagating signals to seemingly arbitrary depth, especially in deep narrow networks where standard PC struggles to learn.
Ultimately, our motivation here is to understand how distributed physical systems, such as the brain, can compute credit signals using only local coupling and local dynamics.
PC-ALM may also inform deep learning in neuromorphic hardware, where dynamics are cheaper than on GPUs.
Paper: https://t.co/doSZ8mzoyK
Code: https://t.co/rxEDIszVKD
I’m excited for this new article in the Notices of the AMS (@amermathsoc). In it, we describe several nice analogies between linear algebra and category theory and discuss connections to language and LLMs. Coauthored with John Terilla and @gastaldi_gianni.
https://t.co/rd3BB0MKOa
Tomorrow @HoldijkLars presents his paper "Path Integral Stochastic Optimal Control for Sampling Transition Paths" (https://t.co/vV5gndJP3R)
I think the SOC ideas might have even more applicability in this field!
Join on Zoom at 11am EDT / 5pm CET: https://t.co/kDchkDIf4U
<not_an_april_fool_joke>
By amplifying human intelligence, AI may cause a new Renaissance, perhaps a new phase of the Enlightenment.
But prophecies of AI doom are also causing a new form of medieval obscurantism.
</not_an_april_fool_joke>
https://t.co/QSwyHdIfGi
ICYMI, I've spent the last several months developing "Gravitas": a new computational relativity tool based on @wolframphysics formalism, enabling anyone to perform hypergraph-based numerical relativity simulations on their laptop. (1/4)
I'm very happy to share that my book has now been printed and will be available next week!
Introduction to Graph Signal Processing: https://t.co/TyofgAb7Il
Numerical Integration to Simulate Nonlinear Dynamics! The final phase of our series on Differential Equations
https://t.co/ahLVeKAeYj
Topics
#Chaos
Forward/backward Euler
Stability & accuracy
Runge-Kutta
Symplectic & variational integrators
Uncertainty propagation
Code examples
Excited to share our generalist neural algorithmic learner! 📚🤖
One graph neural network🕸️, one set of parameters🌐, *thirty* diverse algorithmic tasks 💻 -- delivering on our promise from Oct'21.
It's been quite an amazing journey 🚀
Personal reflections in the thread... 🧵
In our latest work (to appear at two #ICLR2022 workshops!), Andrew Dudzik & I use category theory and abstract algebra to more precisely characterise how GNNs align to dynamic programming.
This further generalises the great work of @KeyuluXu et al!
https://t.co/IOEMlYStJg
Thankful to have found such an insightful conversation between @ilangur and @Ben_Reinhardt.
Shaping Research by Changing Context with Ilan Gur [Idea Machines #36] https://t.co/zNXmI0u7l1
https://t.co/sFTJ8hykMI
New paper (with Xerxes) on the homotopic foundations of @wolframphysics , addressing the fundamental question of *how* and *why* geometrical structures in physics (such as spacetime, Hilbert space, etc.) emerge from discrete "pregeometric" data! (1/5)
Congrats to Ron Dror, @raphaeljlt, Stephan Eismann, and Masha Karelian at @DrorLab and coauthors at @RDasLab for the front page of @ScienceMagazine
Geometric deep learning of RNA structure https://t.co/qbj61aEKmR
Perspective https://t.co/TKQtfTefLT Podcast https://t.co/6MmrfUPZuk
New video on the Sparse Identification of Nonlinear Dynamics (SINDy), 5 years later https://t.co/VoiSrosbtc
Machine learning is enabling the discovery of dynamical systems models and governing equations purely from measurement data.
Our paper demonstrating the power of Bayesian Neural Networks for planetary dynamics comes out in PNAS today!
https://t.co/uc9rIAK3FH (open access)
This paper explores a match made in heaven: chaotic systems and Bayesian neural networks.
Thread:
Tristan Needham's "Visual Complex Analysis" is probably the best piece of technical exposition I've ever read.
I just learned that he published a new book last year, taking a visual approach to differential geometry. I'm so excited to receive my copy!