A Research Scientist at Google DeepMind just dropped a 58 page paper on building agents that specialize in game theory.
Here are the most important parts:
The Top AI Papers of the Week (January 12-18)
- UniversalRAG
- Agent-as-a-Judge
- Self-Evolving Search Agents
- Active Context Compression
- Efficient Lifelong Memory for LLM Agents
- Extending Context by Dropping Positional Embeddings
- Unified Long-Term and Short-Term Memory for LLM Agents
Read on for more:
At MIT, the only course I ever dropped was signal processing. The DFT math was too intimidating. It’s so easy to just type fft() in MATLAB and move on. Years later, I finally did DFT by hand. ✍️ If you are also afraid of DFT, I hope this helps! ⬇️ Download: https://t.co/nNbIF6pn2Z
Amazing continual learning paper out of DeepMind 🚨
Most Continual Learning work assumes the backbone is fixed and the burden on the algorithm to fight catastrophic forgetting. This paper flips that assumption on its head and shows pretty convincingly that architecture choices matter just as much for the plasticity–stability trade-off.
A few takeaways that stood out to me:
Learning vs. retention is heavily architecture-dependent. ResNets and WideResNets are great at picking up new tasks, but they forget aggressively. On the other hand, simple CNNs and even ViTs are surprisingly good at retaining old knowledge, even if they learn new tasks more slowly.
Width beats depth. Making networks wider consistently reduces forgetting and improves average accuracy. Making them deeper often gives diminishing returns on learning while worsening forgetting.
Pooling is a hidden culprit: Global Average Pooling is a major driver of forgetting because it bottlenecks the final representation. Removing GAP or replacing it with smaller pooling layers significantly improves retention.
BatchNorm isn’t always helpful. It helps when task distributions are similar, but under large distribution shifts, BN can accelerate forgetting.
What I’d love to see next is this line of work pushed into the LLM regime (larger models, longer task sequences) so we can (1) benchmark continual learning methods more rigorously and (2) start designing architectures explicitly for continual learning, rather than inheriting them from static training.