AgentSTAR is an agentic method for both reconstruction and tracking of articulated objects from monocular videos.
We can now track through rapid motion, severe occlusion and thin objects! There’s no spoon but we can still track it 🥄
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots.
I wrote a short blog post with my thoughts on the advent of these "robot-use agents."
https://t.co/5bUjZlZxCi
I think it's an important change in the trajectory of robotics!
The transition model is trained on MSLS sequences and evaluated on similarly constrained ground trajectories, so generalization to handheld, pedestrian, or more general 6-DoF motion remains an interesting open question.
TRAIL #ECCV2026 sequential VPR: instead of assuming an ordered reference sequence, it works with an unordered geo-tagged database. Per-frame MegaLoc retrieval is combined with a learned transition score from DINO feature correlations, propagated over time with CRF-style filter.
RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation from Xiaomi EV
One model unifies reconstruction and generation. Given a few posed images and a camera trajectory, it synthesizes a coherent video along the traj.
similar to atlas from worldlabs
Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
tl;dr: fixed 12-frame local context
https://t.co/6RelnH8vPb
The tricky part is 3D Foudantion models can work with only videos. But to add inertial cues such as this paper, one have to do camera–IMU extrinsics calibration .etc which narrows the usage. Besides classic visual inertial odometry or VINS methods works well when calibrated.
VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues
Ernesto Lozano, Alberto Jaenal, @jcivera
tl;dr: IMU + 3D foundation model (3DFM)->metric 3DFM
https://t.co/oB2jXrSuuU
When we say Atlas has pixel-perfect camera control, we mean it.
Atlas precisely follows input camera parameters, including non-planar projections such as the Brown-Conrady distortion model and the Kannala-Brandt fish-eye model.
@BhamidipatiPan1 invented a novel method of camera conditioning and it works beautifully.
🧵 [1/N]