Happy to announce that our paper on Structured State Space Models has been accepted to #AISTATS2026. ☺️
We introduce a geometric framework using warped basis projections to construct Structured SSMs without explicit ODE formulation or discretization.
Interestingly, this formulation naturally leads to a Lag operator interpretation. 📉
Please take a look 👀: https://t.co/KeE1UhFDkf
Traveling waves may underlie working memory (and many other functions).
Traveling brain waves support flexible storage of human working memory
https://t.co/lwA8KpwfnU
#neuroscience
New paper advised by Yann LeCun!
"H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning"
Most world models plan in a single latent space and at one timescale, which makes long-horizon planning expensive and forces one representation to handle both low-level motion and high-level goals.
H-JEPA instead stacks JEPA world models across multiple timescales, with higher levels predicting farther into the future.
The highest level first figures out roughly where the agent should go, then each lower level turns that into increasingly concrete subgoals until the bottom level outputs actual actions.
On Visual AntMaze, this 3-level setup increases success from 18% to 73% while using less planner compute.
https://t.co/dQrd4iT1T9
𝗛𝗲𝘁𝗲𝗿𝗮𝗿𝗰𝗵𝘆 𝗶𝗻 𝘁𝗵𝗲 𝗯𝗿𝗮𝗶𝗻
New preprint with Andrea Gambarotto. Neuroscience often treats control as a hierarchy: the structures behind a behavior hold a stable rank. We argue this mistakes a configuration for an architecture. 🧵
https://t.co/wvvqBPCBWU
Backprop has been the only credit assignment algorithm capable of training large neural nets.
Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!)
- Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers.
- We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization.
- Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones.
- Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling.
The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel.
w/ @bishmdl76, @cs_serdar, @akshayvegesna
Linear attention is arguably the most naive RNN, but still massively outperforms traditional RNNs by maintaining a matrix state. So... what if we use a (triadic) outer product of three vectors and maintain a three-dimensional state?
Introducing: Triadic Linear Attention 🧵
A sweet wrap-up to my visit at @Princeton: my work with @HazanPrinceton and Angelos Assos has been accepted to #NeurIPS2026! 🎉
We show how to learn linear dynamical systems with a small memory footprint. Here’s the idea 🧵
pid control is one of those ideas that looks like an equation until you watch a robot actually move. imagine commanding a robotic arm to follow a trajectory. the controller constantly asks one question: where should the joint be, and where is it right now? the difference is the error. proportional control reacts to the error that exists now. large error means strong correction, small error means gentle correction. increase it too much and the robot becomes aggressive, overshoots and oscillates.
integral control remembers the past. if gravity, friction or some constant disturbance keeps leaving the joint slightly below its target, proportional control may keep producing a persistent error. the integral term accumulates that error over time and keeps increasing the correction until the bias disappears. derivative control looks in the opposite direction: it cares about how quickly the error is changing. if the robot is approaching the target too fast, derivative action pushes against that motion. think of proportional as the spring, integral as the memory, and derivative as the damper.
the interesting part is that pid is really a tiny feedback intelligence loop. observe the system, compare reality with intent, calculate the error, act, observe again. do this hundreds or thousands of times per second and a motor that knows nothing about trajectories can become a precise controlled joint. stack these loops across a robot and you begin turning electrical energy into coordinated physical behavior. before learning exotic control algorithms, understand pid deeply. it teaches the central idea behind control itself: continuously use feedback from reality to correct your model of what should happen.
What's the best way to solve linear regression? (No, it’s not least squares.)
In our new paper, we show that principal component regression (PCR) beats, up to constants, every monotone spectral filter (incl. gradient descent & ridge regression) on every problem instance! (1/8)