GPT-6 Astra is impressive, but agentic robotics ≠ just control.
Physical agents should control robots, learn new skills, and design their own hardware.
So we introduce RLE-Bench: 48 everyday robotics engineering tasks spanning closed-loop control, policy learning, perception, and mechanical design.
“Individual Parameters in Weight-Sparse Transformers Appear Interpretable”
We empirically show that we can describe (with a short Python function!) exactly when a single weight fires (MLP, or attention) in a language model.
Here’s one that fires on the number “100” or more generally “triple digits”: same weight, same job, across totally different sample code inputs. 🧠
New preprint (and Mech Interp workshop paper!)
w/ @sheimersheim 🧵
#MechInterp #weight #interpretability #LLMs #AI #ICML2026 #Mech #Interp #Workshop
It's official! I will be joining @astar_gis as new PI this year. Looking forward to embarking my independent journey in Singapore! My lab will be at the intersection of genomics and RNA technology. We will be hiring at all levels. Email if you are interested! Check below⬇️ #newPI
Excited to share our new work, VeriSpecGen.
A big idea behind the framework is traceable refinement: decompose intent into atomic requirements, generate targeted tests, and use failure attribution for localized repair. This makes spec synthesis much more reliable and data-efficient.
Proud to see strong gains across model families and scales. It was super fun to work on this.
📝 Paper submissions for #CLeaR2026, which will be held on April 6-8 at the @broadinstitute, are due December 15. CLeaR brings together researchers advancing causal discovery, inference, and the role of causality in modern ML. Learn more and submit: https://t.co/GAJIuBO4HQ
In our recent preprint, "The Emergence of Complex Behavior in Large-Scale Ecological Environments," we use DIRT to study how intelligent behaviors emerge in unsupervised populations of agents according to reproduction, mutation, and natural selection.
2/n
NEW in #DeeperLearning blog:
Training quantized LLMs is tough -- the staircase loss landscape kills gradients everywhere.
Meet LOTION🧴: a framework that replaces the quantized loss w/ a smoother objective that preserves the same global minima.
👉https://t.co/BUYWRnyQkv
1/9 Introducing LOTION (Low-precision optimization via stochastic-noise smoothing), a principled alternative to Quantization-Aware Training (QAT) that explicitly smooths the quantized loss surface while preserving all global minima of the true quantized loss. Details below:
Are there conceptual directions in VLMs that transcend modality? Check out our COLM spotlight🔦 paper! We analyze how linear concepts interact with multimodality in VLM embeddings using SAEs
with @Huangyu58589918, @napoolar, @ShamKakade6 and Stephanie Gil
https://t.co/4d9yDIeePd
What precision should we use to train large AI models effectively? Our latest research probes the subtle nature of training instabilities under low precision formats like MXFP8 and ways to mitigate them. Thread 🧵👇
🎉Congrats!
🥇@uthsavc: Mapping the topography of spatial gene expression with interpretable deep learning
🥈Anurendra Kumar: CellWHISPER: Inference of contact-mediated cell-sell signaling
🥉@xinhez: An AI-Cyborg System for Adaptive Intelligent Modulation of Organoid Maturation
Thanks to Jia Liu for hosting a wonderful visit to Harvard SEAS. So great to be back on campus, and see the fascinating work in Allston. Also, shout out to the bright students who were kind enough to share some pizza and a selfie with me.
To all Pennsylvania politicians: I love that @duolingo is headquartered in Pittsburgh and that y'all use it as an example that successful tech companies can start here. If PA makes abortion illegal, we won't be able to attract talent and we'll have to grow our offices elsewhere.
I am still in great shock and can’t believe that Jian Sun passed away today. You probably didn’t know him, but you must have used or at least heard of ResNet, Faster R-CNN, and ShuffleNet among his other highly influential work. It’s a great loss for the entire community! R.I.P.!