We are looking for a postdoctoral researcher in speech and audio processing, with a possible start in the Fall 2026 semester. If you are interested in working with us, please apply through the following form: https://t.co/ssjinOn6LW
📢 ISCSLP 2026 | Penang, Malaysia | Nov 14-17, 2026
The 15th International Symposium on Chinese Spoken Language Processing (ISCSLP) is now open for submissions!
🚀Special Sessions & Challenges Proposals: Jun 8
🚀Paper Submission: Jun 12
🚀Tutorial Proposals: Jun 18
#ISCSLP2026
The 3rd URGENT Challenge just started! 🚀
🎯This time we further expand data diversity and pursue high data quality in the universal SE track.
🎯 I'm also very excited about the new track on SE-oriented SQA. We believe (found) this benefits several areas (SE, TTS, and so on).
Our multi-speaker ASR paper is now available on CSL😊 We integrate SOTA speech separation (TF-GridNet), self-supervised learning (WavLM), and ASR (Conformer-based joint CTC/AED) models in an end-to-end manner. I really appreciate excellent collaborators!
https://t.co/EHJay6hjjo
📢 Introducing VERSA: our new open-source toolkit for speech & audio evaluation!
- 80+ metrics in one unified interface
- Flexible input support
- Distributed evaluation with Slurm
- ESPnet compatible
Check out the details
https://t.co/lKpIaJE6Er
https://t.co/pSzWd5C5YM
Presentation day!
I am going to present our contribution of Team Hamburgers 🍔 to the URGENT Challenge.
NeurIPS, Saturday Dec 14, 3:45 pm, West Meeting Room 215, 216.
@NeurIPSConf
If you are still around in Vancouver for Neurips, tomorrow we will have the URGENT challenge workshop from 1.30 pm.
Come by if you are interested in generalizable speech enhancement (also wind will be up to 70km/h tomorrow and is cozy inside 😉)
Lineup: https://t.co/NCo64Ez8RX
We are thrilled to announce the Interspeech 2025 URGENT Challenge, starting on 11/15!
Join us in building universal speech enhancement models to tackle in-the-wild speech data using large-scale, multilingual data. Details: https://t.co/bZrAqCYwQa
We investigated the scalability of speech enhancement models on different datasets and models (led by @Emrys365 ).
I’ll present this work at #INTERSPEECH2024 Speech Enhancement poster session today at 1:30 pm!
I'm excited to announce @WavLab's XEUS - an SSL speech encoder that covers over 4000+ languages!
XEUS is trained on over 1 million hours of speech. It outperforms both MMS 1B and w2v-BERT v2 2.0 on many tasks.
We're releasing the code, checkpoints, and our 4000+ lang. data! 🧵
Yet another Mamba paper in #INTERSPEECH2024! We explore Mamba's capability in ASR, TTS, SLU, and SUMM on top of ESPnet. Mamba outperforms S4 in the decoder and demonstrates better length generalization and computational efficiency compared to Transformer.
https://t.co/yFd6SfopjO
Excited to co-organize and launch URGENT Challenge as NeurIPS 2024 Competition. The goal of this challenge is to create a “universal speech enhancement” task which goes beyond the usual distortions - additive noise and reverberation. https://t.co/bWkmTylpUD
We are thrilled to announce the URGENT 2024 Challenge - a new speech enhancement (SE) competition at NeurIPS 2024: https://t.co/vXr4jEqeDD
This challenge aims to unify diverse distortions and sampling frequencies using a single universal SE model. #URGENT2024 (1/4)
The validation leaderboard is also live now at https://t.co/g0kArVOuNB. Feel free to participate in our challenge at any time (ddl is Sep 20) and we look forward to your great work! (3/4)
Due to the limited resources, we were unable to cover more architectures (such as Transformer-based, Mamba-based, etc.) and higher complexity (MACs>300 G/s or Param>100M) in our SE scalability investigation. But we do hope our work can ignite more interest in this direction!💪
Ever wondered how speech enhancement (SE) models scale? In https://t.co/ji90K2HQSO (w/ @kohei_1979, @jjungjjee, @chenda54, @shinjiw_at_cmu, and Yanmin Qian), we investigate the scaling effect of several popular SE models w.r.t. dataset size, model complexity, and compute budget.
In terms of model architectures, time-frequency domain dual-path models show great potential in both low-complexity (BSRNN) and high-complexity (TF-GridNet) regions. But they all fail to retain high efficiency when keeping scaling up. We still need to look for a better arch!