Excited to share that I’ll be presenting today our work on “Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models” at the L3D-IVU workshop at @CVPR. Come to poster 125 to find more about our work. (1/4)
We propose a new set of features based on CLIP and CLAP embeddings. Additionally, we propose a simple yet effective model that fuses the information from both CLIP and CLAP text encoders for an enhanced text representation. (3/4)
Excited to share that I'll be presenting our work on "Video-adverb retrieval with compositional adverb-action embeddings" as an Oral today in beautiful Aberdeen at #BMVC2023@BMVCconf! 🎉(1/4)
At #GCPR23 today? Then come to our poster on "Text-to-feature diffusion for audio-visual few-shot learning" by @MerceaOtniel, @hummelth_, A. Sophia Koepke and @zeynepakata! Find out more about our work here: https://t.co/TRUL7tMGLG
Interested in XAI? Do not miss the Explainability in ML workshop!
It will take place March 28-29 in Tübingen (Germany) and will feature talks from prominent researchers in the field.
Check the program here https://t.co/yfjmJhaOE6, and register in advance: few spots remaining!
🗓️Sess 6 | 101, 25/10 | afternoon
“Temporal and cross-modal attention for audio-visual zero-shot learning”
@MerceaOtniel* @hummelth_*, A. S. Koepke, @zeynepakata
We leverage temporal context for better and improved audio-visual zero-shot learning!
Blog https://t.co/9ZcsS3SxJF
If you are at #CVPR2022, come check out one of our six publications!
We have topics ranging from training on partially labelled dataset to (multi-modal) contrastive metric learning and all manners of zero-shot learning!
A quick 🧵 (sorted by date):