Truly a pleasure to be involved in building such a versatile sound event detection model. Huge shout-out to my great collaborators: @apoorv2904@mhnt1580@BowenShi20@bigpon517 👏
Fun example of PE-A-Frame confidently predicting “a man speaking” and “a man coughing”, but being much less confident about “a man predicting the future” 😄
(Though @SchmidhuberAI may have been right back then)
📄 Paper: https://t.co/eF93KFqiwP
💻 Code: https://t.co/xWQXk79mDr
🙏 Many thanks to my supervisor Timo Gerkmann for his invaluable guidance, to the reviewers Shinji Watanabe and Simon Leglaive for their insightful feedback, and to commission members Sören Laue and Jianwei Zhang.
Grateful to all collaborators who made this journey so rewarding!
🎓 I’m thrilled to share that I successfully defended my PhD on generative speech enhancement at the University of Hamburg!
My work explored diffusion-based and audio-visual generative models for robust speech enhancement and audio restoration.
👉 https://t.co/nOkobwi5zE
``Normalize Everything: A Preconditioned Magnitude-Preserving Architecture for Diffusion-Based Speech Enhancement,'' Julius Richter, Danilo de Oliveira, Timo Gerkmann, https://t.co/5TUlsVscqu
Check out the slides here: https://t.co/PNBq4AAyVp. Please note that the PDF is 36 MB in size due to its audio and video content, and it is best viewed using Acrobat Reader.
Join me tomorrow for a webinar on "Generative Audio Restoration in Multimodal Applications"!
I'll introduce the tractable Schrödinger bridge and discuss the differences between flow matching and score-based methods.
https://t.co/hEMANKAFOG
Will present our paper, "Investigating Training Objectives for Generative Speech Enhancement" at #ICASSP2025!
🗓 Wed, 5:00-6:30 PM (AASP-P12)
🔊 Discussing diffusion bridges (Schrödinger bridge) & connections to score-based models—let’s chat!