*NEW PAPER&MODEL*
HOI-DETR for detecting hands, object in-hand, &object-interacted-with (through a tool).
This is the foundation model you've been waiting for! One that works off-the-shelf, single-image, using strong detectors, generalises and is stable
https://t.co/sfzeSq5rZ3
🧵
After two fantastic years at @UCBerkeley I'm thrilled to share that I've joined @Microsoft in Zurich🇨🇭to pioneer the next generation of multimodal foundation models to drive agents 🤖 that can seamlessly interact across the digital and physical worlds 🌍
We are hiring! 🧵
🚀 As #CVPR2025 week kicks off, meet SANSA: Semantically AligNed Segment Anything 2
We turn SAM2 into a semantic few-shot segmenter:
🧠 Unlocks latent semantics in frozen SAM2
✏️ Supports any prompt: fast and scalable annotation
📦 No extra encoders
📎 https://t.co/bdfUd1XPw8
Now on ArXiv our @CVPR#CVPR2025 paper
Learning from Streaming Video with Orthogonal Gradients
Instead of shuffling clips, can we learn from videos fed sequentially, where you see a clip once, in order?
How to deal with the correlation of gradients over training?
1/3
Image segmentation doesn’t have to be rocket science. 🚀
Why build a rocket engine full of bolted-on subsystems when one elegant unit does the job? 💡
That’s what we did for segmentation.
✅ Meet the Encoder-only Mask Transformer (EoMT): https://t.co/LWZXyeR4by (CVPR 2025)
(1/6)
🔥 Our paper SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation is accepted at #CVPR2025! 🎉
We make #SegmentAnything wiser, enabling it to understand text prompts—training only 4.9M parameters! 🧠
💻 Code, models & demo: https://t.co/cUZ8hHOn3Y
Why SAMWISE?👇
🛑📢
HD-EPIC: A Highly-Detailed Egocentric Video Dataset
https://t.co/0y1NQByZdB
https://t.co/0PbngM4oQx
New collected videos
263 annotations/min: recipe, nutrition, actions, sounds, 3D object movement &fixture associations, masks.
26K VQA benchmark to challenge current VLMs
1/N
📢 Our @ACCVConf Oral
It’s Just Another Day: Unique Video Captioning by Discriminitive Prompting
is now on ArXiv
https://t.co/x0krwFBepO
For challenging ego &timeloop movies, uniquely caption ev clip, including those near identical ones, w/out re-training captioning model
1/N
To all our Amigos @eccvconf…
We’re presenting AMEGO this morning. Poster#193
@GGoletto is ready and so am I
Directions: enter poster alley opposite hpc-ai booth, walk to the end, look left… you’ll find us there
Heading to @eccvconf#ECCV2024?
You're on the academic job market for permanent position (Ass Prof - L/SL)?
Already Ass Prof but considering a move?
We're hiring @Bristol University in Computer Vision (Advert out soon).
Ping me (DM or Email) to chat in Milan - AMA about this post
Love timm, but need to do lower-level vision tasks? Meet IMM, a collection of Image Matching Models unified by a simple, easy to use API. You can use any of 30 models (LightGlue, LoFTR, RoMa, SIFT) just by cloning the repo and changing a single parameter! https://t.co/1d65XIMUDo
This work aims to detect and identify activity-centric zones in real-world conditions and leverage their domain-agnostic representations to enhance the generalization of first-person action recognition models.
Paper: Egocentric zone-aware action recognition across environments
Link: https://t.co/QSlpQ5VnoE
Project: https://t.co/2aCG52Kxc5
#AI #AI美女 #LLMs #deeplearning #machinelearning #3D #actionrecognition
AMEGO - a representation of long videos. AMEGO breaks the video into Hand-Object Interaction (HOI) tracklets, and location segments. This forms a semantic-free memory of the video. AMEGO is built in an online fashion, eliminating the need to reprocess past frames.
Paper: AMEGO: Active Memory from long EGOcentric videos
Link: https://t.co/zk0jFkBA5D
Project: https://t.co/Tfh1rkQBVr
#AI #AI美女 #LLMs #deeplearning #machinelearning #3D
📢 [New Preprint]@eccvconf#ECCV2024 paper
AMEGO: Active Memory from long EGOcentric videos
https://t.co/kS6pMIZGKa
Semantic-free representation of all interacting objects and locations in a long egocentric video.
w/ 20K VQA benchmark that uses object crops to query interactions.
Our retrieval + matching pipeline for astronaut photo localization is out! Manually localized photos taken by @Space_Station astronauts are used for disaster management and research - we automate this process with EarthLoc (CVPR24) and EarthMatch (CVPRW24) https://t.co/b2Z4dwpBc0
Waiting @eccvconf reviews - need a distraction?
Check our accepted #IJCV paper
"An Outlook into the Future of Egocentric Vision".
Read about Stanley, our EgoDesigner using #EgoAI to design the set of a movie in 2030, or Marco the factory worker w/ #EgoAI
https://t.co/GWQl66Mecc
Enabling VLMs to understand hand-object interactions in egocentric vision! We use EPIC-Kitchens & @ego4_d to train VLM4HOI.
Check the paper to know how we did it https://t.co/dZGYhjnhKe
Data: https://t.co/xodOJy4WyK
Code: https://t.co/l0WyiLaVv4
Webpage: https://t.co/XgurEHw0rv
A revised version of our paper "An Outlook into the Future of Egocentric Vision" is now available.
This includes the latest peer-reviewed works after our initial submission.
*New* sections on Ego-Language models added and novel insights included
https://t.co/S9lL3REg9E
Applications now open- 2nd Summer of Research @BristolUni#MachineLearning and #ComputerVision (MaVi) group.
R u PhD student with overlapping interests to us?
U can visit for 3mnths this summer!
Apply - DL 19/1/24
https://t.co/geT8PalFOd
Watch 2023 cohort: https://t.co/1q8GGfhhBu
📢"An Outlook into the Future of Egocentric Vision" 44 pages + 385 references survey now available on
@openreviewnet
We invite comments/suggestions/corrections from researchers for 30days. Major contributions will be acknowledged [instructions in 🧵]
https://t.co/qY1ogGToQ1
1/4