🎓PhD Spotlight: Thomas Hummel
A spotlight on our one and only @hummelth_ , who will defend his PhD on 23rd June! 🎉
Thomas started his PhD in 2020 at @uni_tue as part of the IMPRS-IS program under the supervision of @zeynepakata . His main research focuses on multi-modal learning and video understanding using language as guidance, in particular:
📽️👂Tackling audio-visual video classification in challenging zero- and few-shot scenarios
🤸🤔 Enhancing fine-grained understanding of how actions unfold
🎞️💡 Investigating fine-grained temporal reasoning capabilities of VLMs
Beyond this, he also contributed to semantic image synthesis using VQ-models and had the opportunity to intern at @SonyAI_global in Zurich, where his research focused on audio-language alignment for sound effects.
Throughout his PhD journey, Thomas built an impressive publication record—check out the highlights in the thread below. 👇
He was also recognized as an Outstanding Reviewer at ECCV 2024 and CVPR 2024.
We’re incredibly proud of Thomas’s accomplishments and can’t wait to see what he takes on next.
Congratulations on this major milestone! 👏📸🚀
This was a joint work with Otniel-Bogdan Mercea (@MerceaOtniel), A. Sophia Koepke and Zeynep Akata (@zeynepakata). You can find paper, code and more on our project page! https://t.co/hdEcMbIIVO (4/4)
Excited to share that I'll be presenting our work on "Video-adverb retrieval with compositional adverb-action embeddings" as an Oral today in beautiful Aberdeen at #BMVC2023@BMVCconf! 🎉(1/4)
Our method ReGaDa uses a residual gating mechanism to explicitly exploit the compositionality of adverbs and actions when learning text representations. ReGaDa outperforms all prior works on the video-adverb retrieval tasks, setting the new state of the art! (3/4)
Come talk to me at our poster today at #GCPR2023 to dive into the details 🚀
In this work, we propose a novel few-shot audio-visual classification benchmark and a text-to-feature diffusion framework to augment the training!
At #GCPR23 today? Then come to our poster on "Text-to-feature diffusion for audio-visual few-shot learning" by @MerceaOtniel, @hummelth_, A. Sophia Koepke and @zeynepakata! Find out more about our work here: https://t.co/TRUL7tMGLG
Interested in XAI? Do not miss the Explainability in ML workshop!
It will take place March 28-29 in Tübingen (Germany) and will feature talks from prominent researchers in the field.
Check the program here https://t.co/yfjmJhaOE6, and register in advance: few spots remaining!
Come stop by our poster this afternoon (1.B-101) if you want to talk with us about audio-visual ZSL! We show that our temporal and cross-modal constrained attention mechanism outperforms previous work on three audio-visual GZSL benchmarks!
#ECCV2022
🗓️Sess 6 | 101, 25/10 | afternoon
“Temporal and cross-modal attention for audio-visual zero-shot learning”
@MerceaOtniel* @hummelth_*, A. S. Koepke, @zeynepakata
We leverage temporal context for better and improved audio-visual zero-shot learning!
Blog https://t.co/9ZcsS3SxJF