Excited to share that Assembly101 is now hosted on Hugging Face 🤗
https://t.co/nKGRHYTl8U
We've migrated the full dataset and annotations to improve accessibility and ensure long-term availability for the community.
📣 Introducing the Qwen-Robot Suite — Qwen-RobotNav, Qwen-RobotManip, Qwen-RobotWorld, three foundation models, a full stack for embodied intelligence.
🧭 Qwen-RobotNav — the gateway to mobility.
• Unifies 5 navigation tasks in one model: instruction following, point-goal, object-goal, target tracking, autonomous driving
• Controllable observation protocol
• Tool interface for agentic systems
🤖 Qwen-RobotManip — the foundation of interaction.
• Unified state-action space across heterogeneous robots
• Camera-frame delta poses for coherent cross-embodiment training
• Pretrained on a 38,100+ hour open-source corpus
🌍 Qwen-RobotWorld — infinite worlds for physical agents.
• Single world model, 20+ embodiments
• Natural-language action interface
• Predicts physically grounded futures across manipulation, driving, and navigation
Each model is independently useful, and could be composed as physical-world tools.Together, they form the low-level toolkit for general-purpose agentic systems that don't just see the world, but act in it.
📷 Blog:
https://t.co/ytLcbYET26
📖 Report:
Qwen-RobotNav: https://t.co/uPmSwDYGxg
Qwen-RobotManip: https://t.co/GeyIzJSpU8
Qwen-RobotWorld: https://t.co/SXPH1qzDFy
Researchers at RAI Institute developed the Koala gripper to help robots handle tools and objects more naturally.
The system includes a handheld version for humans and a motorized version for robots, both using the same design and sensors.
By learning from human demonstrations recorded with cameras and force sensors, robots can repeat the same tasks using the Koala gripper.
🚨Have work in progress or an accepted @CVPR 2026 paper? Submit to the 2nd VAR Workshop!
🎯Topics include:
• Streaming VLMs
• Real-time activity understanding
• VLM grounding
• Egocentric video understanding
• Language & robot learning
👉https://t.co/jve18jCqUc
Submit to the 2nd VAR Workshop @CVPR :
* Streaming/online vision-language models
* Real-time activity understanding
* Grounding of vision-language models
* Ego-centric video understanding
* Language and robot learning
Details : https://t.co/5n8tfJaC6w
#ICCV2025 decisions have been released!
They are going out in batches to manage server load, so please be patient :)
A huge thank you to our entire community—authors, reviewers, and area chairs—for your hard work.
Congratulations to everyone whose paper was accepted!
One great thing about attending @CVPR from #Europe is the chance to stop over in #Iceland
Roaming this stunning island under the midnight sun, almost completely alone, is amazing ☺️
Come by our poster #135 between 16:00 - 18:00 tomorrow at @CVPR on distilling LLMs for safe and efficient autonomous driving policies.
@deeptibhegde@RajeevYasarla@Matewhs
👉Paper: https://t.co/313sXjq4kb
🚀 Please visit & check out our @angelayao101 @cvml_nus #CVPR2025 paper for condensing datasets of video temporal action segmentation!
🗓️ Today, 17:00 - 19:00
📍 Poster #184
🌐 https://t.co/a2MEd5VA0d
#CVPR2025
Congratulations to @taeinkwon1 and the Team for presenting their work "A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision" as a @CVPR
Highlight #CVPR25 🙌 🎉. https://t.co/HsuGLEljx1
🚀 Check out our #CVPR2025 paper for efficient streaming action detection!
Today @4pm Poster #318.
https://t.co/J7Gee5frQc
✅ SOTA DinoV2 features
✅ Temporal detection metrics
✅ New benchmark w/ diverse domains
🚀 Check out our #CVPR2025 paper for efficient streaming action detection!
Today @4pm Poster #318.
https://t.co/J7Gee5frQc
✅ SOTA DinoV2 features
✅ Temporal detection metrics
✅ New benchmark w/ diverse domains