Glad to share that our work, VGGT-Det (CVPR 2026). Existing multi-view indoor 3D detectors rely on precisely calibrated multi-view camera poses or depth—costly to obtain in real world. VGGT-Det changes this with a Sensor-Geometry-Free framework (no multi-view poses or depth)
This project has been accepted to appear at ICML 2026 😇😇😇 @ICML#ICML2026#ICML.
Our WildActor consistently preserves body identity (including facial features, body shape, and clothing details) under diverse shot compositions, large viewpoint transitions, and substantial motions.
Project page: https://t.co/JttAv6GsJY
Github code: https://t.co/IfOGFXdCWS
arXiv paper: https://t.co/tpzdz8CN2h
Our #ICLR2026 Oral paper MomaGraph will be presented at Oral Session 3D Vision language models II, April 24th 10:30 BRT. Poster session is at Pavilion 3 P3-#1313, April 23rd 3:15-5:45 PM BRT.@furongh will be there, feel free to stop by and chat!💐
We are excited to share our MagicWorld, a video world model that aims to address motion drift and long-horizon error accumulation. It achieves 20 FPS on the L40S GPU and better results on the RealWM120K-Val dataset under VBench. Project page: https://t.co/ha0v0R97Mb
Thanks @HuggingPapers for sharing CARE-Edit! We introduce a condition-aware router that lets heterogeneous experts handle text, mask, and reference-guided edits without stepping on each other. Thrilled to see it accepted at #CVPR2026. Glad to have worked on this w/ @AndyYucheng.
Introducing Nemotron-Terminal: a systematic data engineering pipeline for scaling LLM Terminal Agents.
We bridge the gap between open models and proprietary models with a fully open synthetic-to-real trajectory pipeline.
🤯The payoff: SFT on our Nemotron-Terminal-Corpus boosts Qwen3-32B from 3.4% → 27.4% on Terminal-Bench 2.0 (+24.0), rivaling models multiple its size.
What makes it work?
🌟Terminal-Task-Gen: A lightweight data curation pipeline that seamlessly combines the adaptation of existing datasets with robust synthetic task construction.
🌟Nemotron-Terminal-Corpus: A massive, open-source dataset covering diverse terminal interactions, which contains explicit planning and execution traces for complex long-horizon tasks.
And we’re releasing everything:
📦 Nemotron-Terminal-Corpus (Large-scale dataset)
🤖 Nemotron-Terminal models (8B, 14B, 32B)
Paper: https://t.co/ORIZ01sav1
HF Daily: https://t.co/nSH4hu7I5D
Models & Data: https://t.co/J1Zc22M95r
Our tech report just hit the #1 spot on Hugging Face Daily Papers!
We're also incredibly excited to see the open-source community putting our work to the test, with the Nemotron-Terminal-Corpus dataset currently trending at over 1,800 downloads and counting.
We can't wait to see what the community build with it!
🚨 Deadline extended to March 12 - you still have time to submit a proceedings paper to the OpenSUN3D workshop @CVPR !
If you are working on open-vocabulary 3D scene understanding, 3D perception, or spatial AI, we would love to see your work ✨🤖
➡️ https://t.co/4nSXaJGNpR
On ScanNet and ARKitScenes, VGGT-Det outperforms the strongest SG-Free baseline by +4.4 and +8.6 [email protected], respectively.
Paper:
https://t.co/YN0yGNgzWU
Code (coming soon):
https://t.co/OcsZbYRVXE
Glad to share that our work, VGGT-Det (CVPR 2026). Existing multi-view indoor 3D detectors rely on precisely calibrated multi-view camera poses or depth—costly to obtain in real world. VGGT-Det changes this with a Sensor-Geometry-Free framework (no multi-view poses or depth)
VGGT-Det integrates the VGGT encoder into a transformer-based detector. With Attention-Guided Query Generation and Query-Driven Feature Aggregation, we unlock VGGT’s internal semantic and geometric priors end to end.
The submissions to our @CVPR workshop OpenSUN3D on Open-World 3D Scene Understanding and Representations are open! 🏔️🤖☀️
➡️ The deadline for the proceedings track is on the 5th of March!🏃♀️
🌍: https://t.co/4nSXaJGNpR
📝: https://t.co/TYZKnuKeSo