What if the metadata from acquisition was enough? 👀
Sometimes the most useful labels aren't labels at all.
Introducing📄 Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have
@fedassa Dear @fedassa, I've applied for the internship and would be very much interested in working on scene understanding with VLMs, which aligns with my research: https://t.co/3JbtT1i2sZ
I appreciate your consideration!
Preparing to be🚅@eccvconf ? Take a look at our recent works that we’ll present there.
This year we're happy to share our results covering topics such as forecasting, tracking, domain adaptation, VLMs & LLMs and much more!
Find out more 👇 and come meet us at #ECCV2024
🍕 Heading to Milan for #ECCV2024!
Our team has put together a mega-thread of our papers. I'll be presenting works on LLMs/VLMs, motion forecasting, and corner-case generation for autonomous driving.
Looking forward to great discussions!
🚨@CIIRCCTU researchers @AVobecky and @JosefSivic have developed a new #POP3D method, which was a big success at the @NeurIPSConf 2023.
The work was developed in collaboration with the @valeoai research team, & it's part of the EXA4MIND Industrial Application Case.
Read more⤵️
🚨 Výzkumníci CIIRC @AVobecky a @JosefSivic vyvinuli novou metodu #POP3D, která zaznamenala velký úspěch i na konferenci @NeurIPS 2023. Práce vznikla ve spolupráci s výzkumným týmem @valeoai.
Přečtěte si více zde: https://t.co/0FlZ7Caimc
@EXA4MIND
[@NeurIPSConf'23]🚨Did you miss it? Our POP-3D generates open-vocabulary 3D occupancy predictions from 📷 surround-view images only, w/o human labels & w/ distillation from pre-trained models.
We also propose a new small 3D-occupancy open-vocabulary benchmark #neurips2023⬇️ [1/N]
For open-vocabulary language-driven object retrieval, POP-3D outperforms the MaskCLIP+ baseline (+3.5 mAP) that needs 3D point clouds for projecting image-language features to voxel space. [7/N]
🚨Happy to release on arXiv CLIP-DINOiser: Teaching CLIP a few DINO tricks🦖🎓
We obtain dense CLIP features in 1 forward pass w/o feature alteration and w/ almost no computational extra cost to facilitate open-vocabulary semantic segmentation 🧶
🖥️: https://t.co/vMmMzXpJoc [1/N]
In less than an hour I will present our work 🍾 “POP-3D Open-Vocabulary 3D Occupancy Prediction from Images” at @NeurIPSConf.
We perform open-vocab 3D semantic segmentation w/ only images at inference!
Come to chat at poster #115!
Webpage: https://t.co/MWnzHtYHVs
🍾POP-3D
If you are interested in open-vocabulary segmentation in 3D using only images, come check our poster on Thursday morning!
Poster #115
⏲️Th. 14 Dec. 10:45 a.m. CST — 12:45 p.m. CST
Thanks to @oriane_simeoni D.Hurych @SpyrosGidaris@abursuc@ptrkprz J.Sivic! 🥳