@peteflorence As I was looking through the results, a question came up regarding the y-axis on Figures 1 and 2. The "Error" values differ by about two orders of magnitude between the figures. Could you provide some context on this difference to better align my understanding with your findings?
Thrilled to introduce RoboVLMs 🤖🌍, a unified, open-source and flexible VLA framework for easy integration of any VLMs for robotic tasks, within just 30 lines of code! 🚀 Through 600+ designed experiments, RoboVLMs supports 8 VLM backbones and 4 policy architectures. 📊
What kinds of vision-language-action models exist for embodied AI, and what forms do they take? So far there's been a notable divide between works that focus mostly on vision or vision-language pretraining, robot policy learning, and the text-based "task planning" approaches which mostly predict high-level actions and abstract away motion -- although not always.
From "A Survey on Vision-Language-Action Models for
Embodied AI"
How far can we go with vision alone?
Excited to reveal our Large Vision Model! Trained with 420B tokens, effective scalability, and enabling new avenues in vision tasks! (1/N)
Kudos to @younggeng@Karttikeya_m@_amirbar, @YuilleAlan Trevor Darrell @JitendraMalikCV Alyosha Efros!
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
paper page: https://t.co/oJ4eBQ3gI8
Recent advancements in Large Language Models (LLMs) such as GPT4 have displayed exceptional multi-modal capabilities in following open-ended instructions given images. However, the performance of these models heavily relies on design choices such as network structures, training data, and training strategies, and these choices have not been extensively discussed in the literature, making it difficult to quantify progress in this field. To address this issue, this paper presents a systematic and comprehensive study, quantitatively and qualitatively, on training such models. We implement over 20 variants with controlled settings. Concretely, for network structures, we compare different LLM backbones and model designs. For training data, we investigate the impact of data and sampling strategies. For instructions, we explore the influence of diversified prompts on the instruction-following ability of the trained models. For benchmarks, we contribute the first, to our best knowledge, comprehensive evaluation set including both image and video tasks through crowd-sourcing. Based on our findings, we present Lynx, which performs the most accurate multi-modal understanding while keeping the best multi-modal generation ability compared to existing open-sourced GPT4-style models.
What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?
Presents Lynx, which performs the most accurate multi-modal understanding while keeping the best multi-modal generation ability compared to existing open-sourced GPT4-style models.
proj: https://t.co/TjJNUhy0KQ
abs: https://t.co/sG8hqKYGbR
Great to see the news! HM3D dataset enables us to get 1st place at the Habitat ObjectNav Challenge 2022! Check out the paper here https://t.co/3FBgrxYpab
(1/3) Today we’re releasing the Habitat-Matterport 3D Semantics dataset, the largest public dataset of real-world 3D spaces with dense semantic annotations.
HM3D-Sem is free and available to use with FAIR's Habitat simulator: https://t.co/DbfnjI4X9U
Many thanks to @ylecun for giving a fascinating talk at the Berkeley EECS Colloquium (and for traveling here too!).
Talk archived here:
https://t.co/8o6g208UhB
What target representations make good masked autoencoders? dBOT is a simple, effective, and generalizable masked distillation paradigm. It is out now on arXiv. https://t.co/Hb7CyED4yI