Working on path, motion and mission planning for aerial robots and previously for small ground robots. Also love robot hardware design. Ph.D. Johns Hopkins.
"Contrastive World Models"
If world model has to reconstruct every pixel, it'll pretty much waste most of its capacity modeling irrelevant background noise.
So this paper removes Google DeepMind's Dreamer pixel decoder and instead trains the latent state to identify features of the correct future observation.
It matches Dreamer on clean environments, but performs much better with moving distractors and natural-video backgrounds.
https://t.co/S5b0QWmwuH
Opus 5.5 is so powerful it made my previous workflow for making teaser videos for papers seem over-optimized, which consisted of an iterative loop of storyboarding and rendering.
I simply threw my paper PDFs at it, asked it to generate @3blue1brown-style animations, and let it crunch for an hour. Not only did it generate perfectly coherent animations that adhered to the 3b1b style, it also came up with a lot of nice visualizations and toy experiments that illustrated core concept completely unprompted! Zero hand-holding.
To celebrate entering the last year of my Ph.D., here are the one-shotted animations for some of my selected past projects, starting with Foveated Diffusion (https://t.co/b42nv06I0P).
The moving gaze visualization at 0:16 and the pipeline and 0:24 is so cool!
I asked Opus 5.5 to explain camera focus by building an interactive lens lab
Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost
https://t.co/uqDgUs8H9h
Move the focus ring and you can see the glass elements shift the sharp plane through the scene
I’m back, with a little something…
GPC 1, a new general-purpose classifier enabling: bounding boxes, poses, coordinates, direct mechanical control and so much more, all in milliseconds.
Post-trained on millions of datapoints and thousands of unique problem classes learned directly from you guys posting on X, GPC-1 is a great tool for many things. Early user feedback is also good :)
A new output class enables precise continuous range numerical outputs, so the model can output coordinates of bounding boxes, exact angles and distances, pose estimations, scores, and more. It goes a lot beyond typical classification which just pick a label.
Excited to see where this goes.
Available now!
Opus 5.5 is next level. I’m genuinely impressed by the three.js details.
I generated a playable boat scene through Japanese landscapes with dynamic weather, day and night lighting, realistic textures and 3D characters.
Live site: https://t.co/YMW4Zp52In
The water reflections, physics, scenery and architecture are insane. I could look at this all day.
Two robot arms collaboratively assembling an engine part
Had Opus 5.5 figure out the planning and coordination
I think this is a promising direction for fully automated robotic systems, even if some steps still require more capable hardware
Introducing Real-Time EXPO-FT – Fast and Reliable RL for Real-Time VLA Policies!
Real-Time EXPO-FT unlocks π0.5 on challenging dynamic tasks, such as balancing a ball on a plate and striking a ball into the goal
(1/6)
How far can coding agents go in the physical world — and can what they discover be distilled into fast System 1 robot policies like VLAs?
Introducing EmbodiedSWE, where we systematically study frontier agents across a wide range of precise, dexterous, and long-horizon robot tasks. The results surprised us: these agents could already autonomously solve remarkably complex physical tasks.
But solving is only the beginning. Running a coding agent every time a robot acts is expensive, slow, and difficult to deploy safely in the real world.
So we asked: can we turn what coding agents discover into fast robot policies?
We explore one direction: using coding-agent solutions as teachers for VLAs. A single successful solution can be expanded into a large and diverse training dataset — at surprisingly low cost.
This suggests a compelling path toward general-purpose robots: a GPT-like “brain” for reasoning and discovering solutions, paired with a fast VLA or RL “cerebellum” for acting in the physical world.
For the first time, everyday general-purpose robots don’t feel quite so far away.
Can we steer a robot foundation model without retraining it?
Our new paper explores this question through control-inspired notions of feature observability and controllability in Vision-Language-Action (VLA) models.
We find that state- and action-relevant information can often be linearly predicted from VLA internal representations—and that these representations can be steered at inference time to change robot behavior, without fine-tuning.
As an example, on a real robot, our approach increased the preferred handle-grasp rate from 14% to 74%, while adding only ~1% inference-time overhead.
The broader idea is that representation-level control could provide a lightweight interface for adapting robot foundation models to new preferences and constraints—without retraining.
Joint work with Hugo Buurmeijer Carmen Amo Alonso @SwannAiden
Paper: https://t.co/SMa318wfrt
#PhysicalAI #Robotics #VLA #RobotLearning
@StanfordEng@StanfordAILab
MOSS, the open source litter picking rover currently runs on a Jetson Orin Nano Super and a RealSense. Next, we’re exploring a more affordable build with a Raspberry Pi 5 and cheaper navigation sensors.
Ordered both today:
- Mighty Camera ($92)
- ST’s VL53L9CX 3D ToF evaluation board ($80), with up to 9m range.
The Pi build comes to roughly $820 in estimated parts. Now we need to see how it performs running @dimensionalos stack.
The goal: make MOSS affordable enough for more people to build their own.
Object rearrangement and *reverse rearrangement*! A superhuman capability that's easy for a robot with the right persistent scene representation. Can your "world model" do this?
ReorientBot, @wkentaro, @stephenjames https://t.co/62DPqHkJ7q
MPPI in pure Python, small enough to actually read 🐍
If you've seen Model Predictive Path Integral control cited in off-road autonomy or agile driving work and never sat down with it, this is a good place to start.
You sample a few thousand noisy control sequences, roll each one forward through a model of your system, score them with a cost function, then take an exponentially weighted average of those samples as your next control. Repeat every timestep.
What makes it interestign:
→ No gradients required, so your dynamics don't have to be differentiable and your cost function can be as ugly as you like.
→ No convexity assumption, so it copes with the non-convex problems that give classical MPC trouble
→ Embarrassingly parallel, which is why it took off once GPUs got cheap enough to roll out thousands of trajectories in real time
The repo is lightweight pure Python with scripts and notebooks, and demos covering path tracking, obstacle avoidance, pendulum and cartpole swing-up.
Most production MPPI lives buried inside CUDA kernels and ROS plumbing, and it's genuinely hard to see the algorithm through the infrastructure.
A compact NumPy version shows you the whole idea at once.
Access the repo here: https://t.co/2TNNPLaSvb
~~
♻️ Join the weekly robotics newsletter, and never miss any news → https://t.co/GoA3ZuwoPB
For anyone curious how Jev works, I made a visual explanation using @claudeai :)
This is based on the Qwen2.5-RLCD model which @harshagundal released on @huggingface
The idea is to replace autoregressive LLM generation by a single Transformer decoder (of a pre-trained LLM), which processes the context + JSON schema only once. The keys and values of those tokens are cached.
Next, for each field of the JSON schema, we:
1. pass its field suffix tokens through the Transformer decoder again (reusing the KV-cache)
2. obtain a final hidden state, which we pass through the language modeling head
3. we obtain scores, also called logits, for all tokens in the vocab of the LLM
4. we only look at the scores of the tokens we care about for the given field, and pass those through a softmax to obtain probabilities which sum to 1
5. we take the token with the highest probability.
The benefits of this are that:
1. it's fast (we don't need to generate the JSON schema token by token)
2. it's 100% valid JSON (we don't need to rely on the model to generate a valid schema)
🌍 New World Action Model in LeRobot ✨
LaWAM (CoRL 2026) replaces future-video generation with a single-step latent subgoal.
Given an observation + language instruction, the policy predicts a latent action. A 230M Latent World Model decodes it into future visual features in ONE forward pass, and the action expert conditions on that predicted subgoal to produce the next chunk - dynamics-aware foresight, no iterative video rollout.
🎯 98.6% LIBERO · 91.22% RoboTwin · 90.0% real-world manipulation
⚡ 187ms / action chunk on A100, 10 denoising steps
Thanks for the contribution by Zhongguancun Academy 🙌
Docs: https://t.co/Dc59GaDllL
Blog: https://t.co/ETyQT38nfb
Robots need memory 🧠: remembering where your favorite cup is or which drawer is locked.
How can a robot remember across large spaces, interactions, & fine-grained detail?
Introducing MessyMem (accepted to #CoRL2026 🚀): Learning-from-Doing Memory for Mobile Manipulation. 🧵
World-action models typically imagine the future in RGB – but are pixels really the right representation for robotics?
Our bet is no: RGB spends capacity on fine-grained details and variation that are often irrelevant for robot policies.
DINO features, point tracks, and depth capture more useful features like semantics, motion, and geometry.
But no single modality captures everything — how can we effectively combine them?
We introduce ✨ModAR✨, which predicts the future one modality at a time, with each prediction informing the next. We train from scratch and find that this formulation performs best.
Our 30.1M scratch-trained model even outperforms a 6B video-model-initialized model finetuned on the same data!
��� [1/8]