If you want a vision encoder for dexterous manipulation, what should be the most important part to model? ๐ค
Current standard models like CLIP, SigLIP, and DINOv2 have an incredible grasp of semantics and spatial details. But they lack the action-centric structure needed for downstream visuomotor control.
But collecting annotated robotic trajectories at scale is SUPER expensive and largely unrealistic. So, how do we bridge this gap?
We introduce CAIP (Contrastive Action-Image Pre-training) โฌ๏ธ
๐ธ Action-centric upstream: we align visual observations with action chunks through a contrastive objective.
๐ธ Human video as a proxy: we represent 3D human hand poses analogously to robotic end-effector actions, tapping into a massive source of human demonstrations.
๐ธ Massive scale: pre-trained on over 32,000 hours of manipulation video, driving both sample efficiency and robust generalization.
๐ธ Hardware proven: achieves a 76% average success rate on a real-world Dexmate Vega bimanual @DexmateAI manipulator with dual 22-DoF Sharpa Wave hands @SharpaRobotics .
๐ธ State-of-the-art: significantly outperforms strong baselines like DINOv2, SigLIP, MVP, and Qwen3.5 ViT across complex tasks, even under unexpected lighting changes and visual distractors.
๐ Project: https://t.co/ClUPZvXEDb
๐ Blog: https://t.co/B4wjkUPL0S
๐ Paper: https://t.co/4Xe3bqLJy5
๐ป Model: https://t.co/sGIPPIlmFP
I think the main problem is that we have way too many papers, so when authors have to choose which ones to submit, they'll pick the strongest. That would keep out the half-baked and AI-slop submissions.
In general, charging a fee per submission is a bad idea. It lets well-funded, dominant labs submit as much as they want while labs with less money get squeezed out.
Every few years the field agrees the review system is broken. Then submissions double. ICLR just passed 50,000 ๐คฏ
At some point we have to cap how many papers an author can submit โ๐...Nobody has more than 5 conference-worthy ideas in a single cycle.
@AvivTamar1@ChenTessler And just to be clear, I don't mind listing such folks as authors. But if this person is limited by the 5-paper cap, something is off. It doesn't make sense to me that someone contributing at that level would have more than five papers.
@AvivTamar1@ChenTessler I barely work with students who contribute less than two full days a week on a project. Of course, there can be a more senior person who does less work than the others โ but if someone isn't contributing much, they should be acknowledged rather than listed as an author.
@ChenTessler Time doesn't equal contribution, but I've never seen someone put in 2โ3 hours a week and make a significant contribution โ I can't think of a single case.
Obviously there could be an exception, but I doubt a student like that would be blocked by the 5-paper cap, right?
@ChenTessler Yes, of course. An acknowledgment is perfectly fine.
Think about it the other way around: working 8 hours per week across 8 different projects DOES NOT justify authorship on eight submissions. Right?
I've been playing with @Muse recently, but still haven't figured out the best way to use it.
What are the good use cases people are actually using it for?
For example, it refuses to control my WhatsApp and send automatic messages, but it sends on Instagram, which is really not useful.
@VarunGangal Even for senior researchers, it shouldn't be more than 5โ6 papers. At the end of the day, there's a limit to how many papers you can actually contribute to.
And I'm saying this from both sides of the table.
@selfattentive Even for senior researchers, it shouldn't be more than 5โ6 papers. At the end of the day, there's a limit to how many papers you can actually contribute to.
And I'm saying this from both sides of the table.