🚀 Happy to share our #INTERSPEECH2025 paper:
Using speaker & acoustic context, we dynamically adjust model paths, resulting in a 25.7% relative BLEU improvement in speech translation.
We also analyze how context influences model behavior.
📜 Paper: https://t.co/t6VQ1JxJdS
Compiled list of @WavLab's first-authored papers (led by @shinjiw_at_cmu) at #INTERSPEECH2024
Wide range of topics, including weakly-supervised/self-supervised speech models, discrete-unit ASR, speaker embeddings/verification, singing voice, blind source separation, etc: 👇
What an honor to be in the cockpit while researchers from CMU, Fudan University, UC Berkeley and NVIDIA developed the approach what won DCASES's 2024 Audio-to-Text Captioning challenge!
https://t.co/1jJ2bAa7Kh