Introducing Meta Segment Anything Model 2 (SAM 2) — the first unified model for real-time, promptable object segmentation in images & videos.
SAM 2 is available today under Apache 2.0 so that anyone can use it to build their own experiences
Details ➡️ https://t.co/eTTDpxI60h
🎉 Excited to invite you to our poster presentation of GLaMM: Pixel Grounding Large Multimodal Mode tomorrow at #CVPR2024 at 10:30 at poster #326🔥
1️⃣ Tackles multiple vision-language tasks in one model.
2️⃣ Fully automated GranD dataset with 7.5M concepts.
3️⃣ Fully open-source.
⚡How can complimentary strengths of image & video encoders help Video-LMMs? We present new results + diverse instruction data + benchmark!
🔗 VideoGPT+ : https://t.co/EM5ns1uKn0
📈Exciting updates to our recent effort to extend LLaMA3 and Phi3 for *visual* understanding. Enjoy!
💻Online demo: https://t.co/yHcQ8niYmd
📓 Chat in Google Colab: https://t.co/C4ULZzSSnd
🚀LoRA, fully FT and S2 FT models added! https://t.co/o4AVEn0AYF
GLaMM: Pixel Grounding Large Multimodal Model
paper page: https://t.co/HKZxWVSsOx
Large Multimodal Models (LMMs) extend Large Language Models to the vision domain. Initial efforts towards LMMs used holistic images and text prompts to generate ungrounded textual responses. Very recently, region-level LMMs have been used to generate visually grounded responses. However, they are limited to only referring a single object category at a time, require users to specify the regions in inputs, or cannot offer dense pixel-wise object grounding. In this work, we present Grounding LMM (GLaMM), the first model that can generate natural language responses seamlessly intertwined with corresponding object segmentation masks. GLaMM not only grounds objects appearing in the conversations but is flexible enough to accept both textual and optional visual prompts (region of interest) as input. This empowers users to interact with the model at various levels of granularity, both in textual and visual domains. Due to the lack of standard benchmarks for the novel setting of generating visually grounded detailed conversations, we introduce a comprehensive evaluation protocol with our curated grounded conversations. Our proposed Grounded Conversation Generation (GCG) task requires densely grounded concepts in natural scenes at a large-scale. To this end, we propose a densely annotated Grounding-anything Dataset (GranD) using our proposed automated annotation pipeline that encompasses 7.5M unique concepts grounded in a total of 810M regions available with segmentation masks. Besides GCG, GLaMM also performs effectively on several downstream tasks e.g., referring expression segmentation, image and region-level captioning and vision-language conversations.
🚀 Exciting times ahead at #CVPR2023!
I'm thrilled to share that our paper, "Fine-tuned CLIP models are efficient video learners" (aka ViFi-CLIP), has been chosen for presentation. Catch us at Poster #231 at CVPR 2023! Can't wait to see you there. 🎉
- Hanoona Bangalath 🇮🇳
https://t.co/qYFofOxahr. in Computer Vision
“MBZUAI supported both my professional & personal growth. Soon after joining, we were introduced to state-of-the-art research &
were motivated to delve into cutting-edge studies.
#MBZUAI2022#MBZUAIGRADS
I am excited to announce that our paper "Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection" has been accepted at #NeurIPS22.
You are welcome to attend our poster session on Nov 30, 09:00 AM - 11:00 AM (PST) @ Hall J (137).