🚀 Introducing BAAI WuJie·RoboBrain Orca — an early step toward Multimodal Latent World Models.
🔥 Instead of predicting only the next-token, next-frame, or next-action prediction, Orca learns world latent representations from multimodal data, and models how world state transitions, i.e., "Next-State-Prediction" modeling.
❄️ After learning the representations, the Orca backbone is frozen, and only a lightweight decoder is post-trained to readout text, images, and actions.
🌍 As the pretrained data scales, latent ability gradually improves, and downstream performance continues to increase. Furthermore, Orca outperforms baselines with similar scale across different tasks. Particularly in embodied tasks, Orca trains the action expert from scratch using only 200 robotic data, and its results in OOD scenarios all surpass those of large-scale pre-trained pi0.5 by binary success rate, white-box rule-based scoring, and black-box evaluation. It provides a new way to address the problem of poor generalization caused by limited robotic data.
🌟 Orca's goal is not to claim “the strongest”, but to share a scalable Next State Prediction paradigm, along with the insights, limitations, and open questions the BAAI Orca Team found along the way.
Paper: https://t.co/7ez8HPSpW2 Project: https://t.co/AjRnERaje2
(7/n)Lego leverages egocentric vLLM to generate enriched action description, and feed the enriched description together with the vLLM embeddings to SDMl to generate an action frame vividly depicts how an action should be conducted based on the user’s visual context.
(5/n)The rest of the story is obvious. By marrying egocentric vision and generative models including vLLM and SDM, there is an exciting opportunity to directly generate a pixel-level representation to address the howto problem.
(2/n)Baby step, we started with action recognition. Yet, action recognition will serve as a trigger of all the magical things we envision the egocentric AI system will offer. For the howto problem, the model needs to output fine grained representation than simply answering whatis
(4/n)After joining Meta GenAI, I start to realize that the foundational models incorporate the knowledge of human skill, but still need customization so that they can be applied to the user’s current situation. What more straightforward way then egocentric visual perception?
(3/n)we start to look at gaze, hand location/mask, 3D body, 3D scene… Yet, these representatives are still not ideal for skill transfer, as users would want something can be easily interpreted, like the instructions from the LEGO, but with more customized to user setting.
(1/n)When I first started my PhD with Jim, one thing we keep talking about is leveraging egocentric vision for skill transfer, to help robot or human to solve the howto problem.
Our paper was awarded the Best Student Paper Prize in BMVC2022🎉 Thanks for my advisor @RehgJim and all co-authors @aptx4869ml@fionakryan. Now we have released our data, codes and pretrained weights on GitHub (https://t.co/Kl2iESLPN4) as well as a video demo on the project page.
Our paper was awarded the Best Student Paper Prize in BMVC2022🎉 Thanks for my advisor @RehgJim and all co-authors @aptx4869ml@fionakryan. Now we have released our data, codes and pretrained weights on GitHub (https://t.co/Kl2iESLPN4) as well as a video demo on the project page.
How can we fill in missing pulsative sensor data? Prior state-of-the-art fails in our novel setting, despite its well-defined temporal structure.
Checkout our #NeurIPS2022 paper, PulseImpute, @ 4 pm CST!
arxiv: https://t.co/Hbv2x7ZvkP
github: https://t.co/bTManuLyEH
Dense self-supervised learning from multiple 3D viewpoints → dense feature representations that generalize both to novel object instances and to novel categories of instances.
Checkout our #NeurIPS2022 paper!
arxiv: https://t.co/rxzdrScII1
github: https://t.co/5n98W7Wykt
Our team at RLR is hosting a @CVPR tutorial for always-on egocentric vision research using Project Aria on this coming Sunday afternoon. Together with tutorial, we also released the first Project Aria Pilot Dataset and data tools.
Tutorial page: https://t.co/gRnn0gF0C6
Join the #Ego4D challenge, exploring the largest ever dataset of first-person video and five new research benchmarks: episodic memory, hands+objects, social, AV, forecasting. First round of the competition ends June 1 with results shared at #CVPR. https://t.co/cmL7qmfzKH