We are presenting WFM-Eval at two @CVPR 2026 workshops in Denver π
ποΈ Jun 3, Video World Models
Poster 9:50β10:40 AM, Exhibit Hall A
ποΈ Jun 4, Foundation Models Meet Embodied Agents
Poster 3:55β4:30 PM
Come say hi π
Work done with @AmberZhang99@prithvijitch@judyfhoffman
How can we run reconstruction models like ΟΒ³ and Depth Anything 3 in real-time?
We present KV-Tracker, a training-free approach, for real-time tracking of scenes and objects. Achieving up to 30 FPS!
With @alzugarayign, @makezur, @XinKong_IC and @AjdDavison
3. Bolt3D (ICCV 2025): Multi-view diffusion model that jointly generates color and 3D scene coordinates
https://t.co/7XUpM43J96
4. World-consistent Video Diffusion (CVPR 2025): Video DiT that jointly generates both color and 3D scene coordinates
https://t.co/yIE28hWKqR
It was fun presenting our work on "Orchid" - a Joint Color and Geometry Diffusion model at ICCV this year. πΈ
Orchid is a single model we trained to generate color, depth, and normals from a text prompt, all at once.
π§΅ with a summary and some connections to concurrent work
1. JointDiT, (also at ICCV 2025): joint color and depth generation (no normals). Unlike Orchid, they use separate latents for color and depth.
https://t.co/BSflPJf696
2. MMGen: color, depth, normals, and also segmentations (for ImageNet classes)
https://t.co/bBP9AVhaLv
New paper π’π’
TL;DR: We introduce OmniNOCS: a multi-domain NOCS dataset with 90+ object classes. Our model trained on OmniNOCS can predict 6DoF object pose and partial shape from their 2D detections, even generalizing to in-the-wild images!
Background & details in π§΅ below
New paper π’π’
TL;DR: We introduce OmniNOCS: a multi-domain NOCS dataset with 90+ object classes. Our model trained on OmniNOCS can predict 6DoF object pose and partial shape from their 2D detections, even generalizing to in-the-wild images!
Background & details in π§΅ below
New paper π’π’
TL;DR: We introduce OmniNOCS: a multi-domain NOCS dataset with 90+ object classes. Our model trained on OmniNOCS can predict 6DoF object pose and partial shape from their 2D detections, even generalizing to in-the-wild images!
Background & details in π§΅ below
OmniNOCS has been accepted to ECCV 24! π₯³
Please consider using it to train or evaluate models for 3D object coordinate regression.
Joint work with my awesome collaborators @_abhijit_kundu_, @kmaninis, Matthew Brown at @GoogleDeepMind and @jhhays at @GeorgiaTech@mlatgt .
We trained a new model (NOCSformer) on OmniNOCS to predict the NOCS and pose of objects from RGB alone (no depth input!). The large-scale training allows it to generalize to in-the-wild images, including objects from COCO.
Data, paper, more results at: https://t.co/5mDKqH3CFO