new blogpost about beginner-level advice on breaking down proofs, which I feel students new to theory may find helpful especially at this time where most derivations are done with the help of LLMs: https://t.co/CHB7eInm0N
It’s that time of year again: graduate admissions season.
Inspired by @phillip_isola’s post today on doing a PhD in the age of AI (highly recommended!), I share some thoughts on how to email a professor about joining their lab.
It’s a first cut, one person’s perspective, and meant to be a living document and discussion starter. Add your perspective via Margin Notes!
We check OOD generalization for @periodiclabs Neon and report optimal cost-performance here, too. This indicates useful generalization for our model.
But our experimental data also opens up an interesting ML research program.
For example, RL tasks to directly predict experimental outcomes given all prior experimental data known up until that date. This is loosely analogous to the often-proposed ML experiment of training AI on all literature before 1905 and then evaluating whether it can derive the theory of relativity.
An AI that already knows the answer through pre-training can cheat on RL tasks like this (i.e., it can skip reasoning and simply output a memorized experimental result). A unique, complete, dated system of record makes this work possible and fruitful.
What if we explicitly learn shorter-horizon values before the longer-horizon values that depend on them?
DCRL builds on @seohong_park’s Transitive RL and its divide-and-conquer view of value learning. This simple change turns out to matter a lot.
More details from @yeonsumia 👇
I'm reflecting on how much research has changed since I've joined the PhD and wrote a short blog post about it (I joined in the tail end of the BERT era!). It seems pretty crazy how different processes are now, and I took the chance to do a retrospective before graduation:
https://t.co/OR7N8s4CgO
Behavioral cloning mystery
https://t.co/VqxzvcfGSx
I wrote a new blog post about "mysteries" in behavioral cloning that appear with real-world robot data (e.g., overfitting is "good"). I also tried to demystify them and shared my thoughts!
Action chunking is a mysteriously effective method. Modern large-scale imitation learning basically doesn't work without it. But why does it actually help? In our new paper, we try to break down the reasons. As the saying goes, what happened next might surprise you...
Meta learning and recursive self-improvement are old ideas. Foundation models breathe new life into them. Our new survey, “Self-Improvements in Modern Agentic Systems,” reviews how the concepts are continuing to evolve.
Paper: https://t.co/59oXCMVUkD
Project: https://t.co/sZwFYdGetH
Github: https://t.co/7OFgJUCN3a
New blog post:
https://t.co/T29LEH7lpY
This is the second in a "series" of posts about how to define and work on research problems. (The first post was 3 years ago...)
This one addresses the agentic elephant in the room.
Some personal news: I've joined Google DeepMind as a Research Engineer on the Robotics team, just wrapped up my first week! So grateful to be part of such a thoughtful team pushing the frontier of embodied intelligence.
Now that I'm settling into the Bay, I'm looking to start or join a group house in the South Bay or SF. I'd love to live with people who are equally obsessed with frontier AI/robotics, and create a space together where we can connect over dinner and talk about anything from dexterity to world models, or just life.
Flexible on timing – anywhere from Aug 1 to Sep 1 works. If you're interested or know of a house, DMs are open!
We just got the top community score on ARC-AGI-3's 25 public games: 78.4% and 160/183 levels cleared, where the best frontier LLM session scores 7.8%. No training or demonstrations allowed.
Paper: https://t.co/mKAVXayoQ6
Blog: https://t.co/1i5pPyGxS7
w/ @WenhaoLi29@ScottSanner
#AI #worldmodel #agentic #AGI @arcprize@fchollet
Anthropic pays $750,000+ a year for engineers who can build LLM architectures from scratch. Stanford taught the entire thing in 1 hour lecture & released it for free.
Bookmark & watch this today before someone takes it down ...
This week's #PaperILike is "Hypothesis-driven Model Expansion under Uncertainty for Open-World Robot Planning" (Xiao et al., RSS 2026).
Really impressive open-world long-horizon mobile manipulation examples: https://t.co/lf7MqVy9iv
PDF: https://t.co/OHiJQ2GdHV
🧠Model-Based RL shows promises but has seen limited success in real-world robotics.
🌎Introducing Robotic World Model, a black-box end-to-end neural dynamics model that bridges this gap, where policies are trained purely in imagination.
@NeurIPSConf
🎯https://t.co/6YYsiVWZic
Here's another paper we'll be presenting at ICML'26 - @icmlconf.
Laplacian Representations for Decision-Time Planning
https://t.co/0UEIuDxS47
This work was led by @DikshantShehmar, and he'll be presenting on 8th July at 2 pm, Korea time.
This video summarizes the idea well ↓
World models are increasingly central to how agents learn and plan.
Today we're releasing WorldModelGym, a benchmark built around a single question: if an agent uses a world model to choose among actions, does it pick the right one?
We call this decision-based fidelity. 100+ tracks across Atari, Meta-World, DeepMind Control, and classic control. One frozen policy. Reality scores it.
Read the full post → https://t.co/OzVd1n6Vth
As a PhD student, I was told not to work on deep RL - too full of hacks and alchemy" But after a year or two of working in this area, I’ve come to (deeply?) appreciate all of the thoughtful research that’s gone into understanding what/why things work and how to make them better.
My lab (joint with @abhishekunique7) took what we’ve learned by reading this body of literature to answer the question: what are the actual best-practices for finetuning a diffusion/flow/generative robot policy (for now, in sim)? Under one set of constraints - ample compute but limited time on your robot - @servo97 paper gives a pretty compelling answer.
Temporal Difference Learning for Diffusion Models (ICML26) https://t.co/pXwzmG3Apm
By Yangchen Pan (my former PhD student) and co.
It reformulates diffusion training as a Markov reward process and introduces a TD obj to encourage temporal consistency across denoising steps.