[1/N] Large flow policies (e.g., VLAs) have become extremely popular, unlocking scalability. However, they come at a cost: the robot cannot react quickly. π𝐑² enables them to achieve reactive, real-time control.
[1/N] Introducing Point4D: dense 4D reconstruction across hundreds of frames from a monocular video. Unlike existing feed-forward 4D methods that operate over short windows, Point4D keeps tracking in 3D, even across occlusions and out-of-frame motion.
My group at Johns Hopkins now has a name and a website
Introducing the Brains, Bots, and Behavior Lab
https://t.co/Iokm3YxuXQ
We're always looking for creative researchers at all levels to join us!
this policy reacts quickly while still using a VLA.
two insights:
1. two channels - "slow" async channel for images, language, "fast" sync channel for touch, torque, joint state at every denoise step
2. sliding window - sliding noise schedule avoids full recalc
Congrats @sungj1026!
Large-scale flow matching policies have low reactivity (open-loop action execution within a chunk) and high perception-to-action latency (VLM inference + multi-step diffusion inference). We show how to make them useful for real-time reactive control.
https://t.co/BxysfzL7AT
[1/N] Large flow policies (e.g., VLAs) have become extremely popular, unlocking scalability. However, they come at a cost: the robot cannot react quickly. π𝐑² enables them to achieve reactive, real-time control.
Introducing MotionForesight:
A very simple approach re-purposing video models for forecasting future 3D scene flow given a short observed video context
Trained with just 40k RGB human videos, it generalizes well to completely new home/office scenes!
1/n
RayRoPE is accepted to ECCV '26!
Need to condition your model on camera poses? Give RayRoPE a try!
In our latest version, we add more experiments and show that RayRoPE can be a plug-and-play module that improves performance in various 3D tasks!
I will be at ICML next week to present this poster.
Happy to meet and discuss with everyone there! If you are interested in discussing about generative modeling/ world modeling, please feel free to DM me.
To be honest, training on handmade 4D asset datasets is a dead-end. Almost all 4D asset data is synthetic and diverse real data barely exists, so models trained on it struggle to reconstruct objects that deform, get occluded, and move freely about the scene.
Our new work, Lift4D, instead lifts 2D & 3D priors into 4D, reconstructing complete dynamic objects from a single in-the-wild video 🧵 (1/n)
🔗Webpage + Demos: https://t.co/XI5jUViTpC
💥Introducing FACTR 2, learning external force sensing on commodity robot arms without needing dedicated sensors.
We show that learned force signals enable force-feedback teleop on low-cost arms and improve BC policies.
FACTR 2 consists of:
1. Neural External Torque (NEXT): learns external forces without needing dedicated force sensors.
2. Force-Informed Re-Sampling Training (FIRST): uses the learned force signal to identify task-critical regions and upsample them during training.
w/ @StevenOh_@_tonytao_
🧵(1/N)
Attending #ICRA2026 in Vienna this week! Feel free to reach out if you wanna chat.
Thurs: DemoDiffusion (https://t.co/Z37ZJ6bInk) at main poster 9-10:30 am.
Fri: DemoDiffusion, Dex4D(https://t.co/vRZehvRdg6) at Beyond Teleop workshop.