👋Excited to introduce PixelRefer, which excels at understanding any object in images and videos:
- PixelRefer-7B outperforms DAM-8B and PAM-3B on both image (e.g., DLC-Bench) and video benchmarks (e.g., VideoRefer-Bench).
- PixelRefer-Lite-2B is ~3x faster than DAM-3B with ~2x less memory footprint.
- Achieving competitive results using only 2.2M data samples, significantly less than other methods.
All the code, models, and data are already available:
Homepage: https://t.co/yhQFO8ebKT
Demo: https://t.co/BNaPE7JVdF
Code: https://t.co/DA5rPZNwe5
HuggingFace: https://t.co/YKCnXqmcRO
Introducing RynnEC, our little step towards physical world understanding🚀🚀🚀
1. RynnEC is object-centric, supporting the recognition of up to 12 object properties/relations.
2. RynnEC is space-aware using RGB videos only (45.8 on vsi-bench), no explicit 3D encoding required
3. Most importantly, RynnEC is able to map language queries to semantic masks, lower ambiguity and easier to be integrated into downstream embodied agent/policy!!!
Project page: https://t.co/jvIhY5RLOB
Blog: https://t.co/Ft8qaMXWOk
#CVPR2025
Announcing VideoRefer x VideoLLaMA3!
🚀🚀Large performance leap from the original VideoRefer-7B
💪💪Region-level vision-language understanding for not just video but also image inputs
✊✊ Surpassing DAM-8B and PAM-3B on both image and video benchmarks
Demo: https://t.co/hvYbZTZXm1
Code: https://t.co/qdJrSRZyuP
Weights: https://t.co/d73tFnwJKG
#CVPR2025 Picks #3
Alibaba just released VideoRefer-VideoLLaMA3 (2B & 7B video LLMs with A2.0 license!)
These models can understand videos and segment objects, answer questions about them throughout the video at the same time 🤯
see it in action ⤵️
Introducing VideoRefer Suite: the key to precise video object understanding! 🔑
Understand any objects you're interested within a video! 🧐
Dive into our model, dataset, and benchmark via the links below! 🙌
GitHub: https://t.co/7lc1lHWjPm
Osprey: Pixel Understanding with Visual Instruction Tuning
Understand everything for SAM! 🙌We introduce a pixel-level multimodal language model Osprey, please check out our paper and demo.
paper page: https://t.co/YNaz1nKBli
code link: https://t.co/case1tLieJ
We released Osprey-724K🥳, an instruction dataset with mask-text pairs, containing around 724K GPT-generated multimodal dialogues to encourage MLLMs for fine-grained pixel-level image understanding.
🎁dataset: https://t.co/htZ8pqcWZx
🔑code: https://t.co/case1tLieJ
We propose Osprey, a mask-text instruction tuning approach, to extend MLLMs by incorporating fine-grained mask regions into language instruction, aiming at achieving pixel-wise visual understanding.
Delighted to share our latest work - “Osprey: Pixel Understanding with Visual Instruction Tuning”🥳
arxiv preprint: https://t.co/jd5Mviy8CB
code: https://t.co/case1tLieJ
video demo: https://t.co/51ta3ZvuFC
Online demo can be found in our github. Welcome to try it!🤩