Thanks for sharing!
I think these are great examples to showcase generalization to wide concepts; smoke and cracks are quite challenging, and likely cases where finetuned methods would fail. I would say mask quality being not exceptional is expected, but I think still appreciable for a feature-only method (curious of which resolution is used, and if CRF is applied). Overall, I think the biggest challenge is identifying the correct concepts. If compute is not an issue, the tradeoff can be resolved using SAM for mask refinement (derive box and prompt).
✨ #CVPR2026 Oral ✨
INSID3 turns a frozen DINOv3 into a training-free in-context segmenter across domains and granularities!
Excited to present our work today in Oral Session 4D (14:00–15:15). Come by our poster later if you’d like to chat:
📍 Poster #19
🕓 16:00–18:00
See you there!
Just landed in Denver for #CVPR2026! 🏔
Excited to present two oral papers:
📍 June 6, 14:00–15:15 — Oral Session 4D: Visual Segmentation
• INSID3: Training-Free In-Context Segmentation with DINOv3
• MARCO: Navigating the Unseen Space of Semantic Correspondence
Posters later that day, 16–18:
• #19 — INSID3
• #20 — MARCO
Also on June 4:
• MARCO @ Workshop on Image Matching
• INSID3 @ Workshop on Visual Concepts
And don’t miss our demo:
🗓 June 5, 16–18 — FoundYou
Happy to chat, come find us around! 👋
First stroll in the poster area at @CVPR... It appears the board length is too short for the advertised poster length (the usual 2m). If the boards aren't going to be replaced, poster sessions of the main conf will be quite a mess. #CVPR2026
Excited to be at #CVPR2026! 🌟
I’ll be presenting two Oral papers: INSID3 and MARCO
📍 June 6, Oral Session 4D
Orals: 14:00–15:15
Posters: 16:00–18:00, #19–20
Also around CVPR:
• June 4: MARCO @ Workshop on Image Matching
• June 4: INSID3 @ Workshop on Visual Concepts
• June 5, 16:00–18:00: FoundYou demo
I’m also exploring job opportunities from next year, DM me if you want to chat! 👋
3) To evaluate beyond standard benchmarks, we introduce a new benchmark with 62 novel categories from diverse domains, including animal species, furniture, and clothing, all unseen during training.
We also include novel keypoint annotations on known categories, increasing coverage (e.g., from 7 to 68 landmarks for human face)
This enables evaluation along two axes of generalization: to novel categories and to novel keypoints within known ones.
What if a model could learn dense semantic matches from just a handful of annotated landmarks, while still generalizing to unseen keypoints and categories — and running 10× faster than diffusion-based approaches?
MARCO is selected as an Oral at #CVPR2026! A unified model for generalizable semantic correspondence, built on DINOv2⭐️
👉 Try our model: https://t.co/Zvt4QTRVJQ
✨#CVPR2026 Oral ✨
A tale of a failed experiment: what if you fine-tune DINOv2 on sparse keypoints, beat every benchmark, only to discover it performs worse than the original frozen model on novel keypoints?
🚀MARCO closes this gap: a unified model for generalisable correspondences
https://t.co/vE62YiTVfd
What if a model could learn dense semantic matches from just a handful of annotated landmarks, while still generalizing to unseen keypoints and categories — and running 10× faster than diffusion-based approaches?
MARCO is selected as an Oral at #CVPR2026! A unified model for generalizable semantic correspondence, built on DINOv2⭐️
👉 Try our model: https://t.co/Zvt4QTRVJQ
✨ As a first-year PhD student, I used to wonder what it must feel like to have a paper selected as an Oral at #CVPR. Today, I’m experiencing that feeling twice!
I’m beyond happy to share that both of my first-author papers have been selected as #Oral at #CVPR2026 🎉
🔥 Can in-context segmentation emerge directly from frozen DINOv3 features?
At #CVPR2026, we present INSID3: Training-Free In-Context Segmentation with DINQv3 — a collaboration between PoliTo, TU Darmstadt and TU Munich.
A training free approach that generalizes from object-level to part-level and personalized segmentation, across natural, medical, underwater, and aerial domains
Check it out: https://t.co/AaMRbfjLyn
@gabTrivv@Meta@p_bojanowski Congratulations on this wonderful achievement! It’s a great joy and pride to see where your talent, passion, and dedication have taken you. Wishing you all the best for this exciting new journey at Meta
I can officially say #PhDone! 🎓
This chapter has been an amazing learning experience. Along the way, I learned that papers mattered less than the people, the late-night experiments, the lively conference conversations.
A huge thanks to all the people that made this possible👇
🚨CVPR 2025 Highlight Paper Alert 🚨
➡️Paper Title: SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
🌟Few pointers from the paper
🎯Referring Video Object Segmentation (RVOS) relies on natural language expressions to segment an object in a video clip.
🎯Existing methods restrict reasoning either to independent short clips, losing global context, or process the entire video offline, impairing their application in a streaming fashion.
🎯In this work, authors aimed to surpass these limitations and design an RVOS method capable of effectively operating in streaming-like scenarios while retaining contextual information from past frames.
🎯They build upon the Segment-Anything 2 (SAM2) model, that provides robust segmentation and tracking capabilities and is naturally suited for streaming processing.
🎯They made SAM2 wiser, by empowering it with natural language understanding and explicit temporal modeling at the feature extraction stage, without fine-tuning its weights, and without outsourcing modality interaction to external models.
🎯To this end, they introduced a novel adapter module that injects temporal information and multi-modal cues in the feature extraction process.
🎯They further revealed the phenomenon of tracking bias in SAM2 and proposed a learnable module to adjust its tracking focus when the current frame features suggest a new object more aligned with the caption.
🎯Their proposed method, “SAMWISE”, achieves state-of-the-art across various benchmarks, by adding a negligible overhead of less than 5 M parameters.
🏢Organization: Politecnico di Torino [@PoliTOnews ], @FocoosAI
🧙Paper Authors: Claudia Cuttano, @gabTrivv , Gabriele Rosi, @masone_carlo , Giuseppe Averta
📝 Read the Full Paper here: https://t.co/boAuqTURIr
🗂️ Project Page: https://t.co/U1OmhvjtDp
🧑💻 Code: https://t.co/HWs5683DeM
🎥 Be sure to watch the attached Demo Video - Sound on 🔊🔊
🎵 Music by Adi Iswanto from @pixabay
Find this Valuable 💎 ?
♻️QT and teach your network something new
Follow me 👣, @NaveenManwani17 , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements.
#CVPR2025 #highlight