Our short paper has been accepted to EMNLP 2025 Findings! We proposed an efficient/effective pipeline to analyze and mitigate social biases in large-scale pretraining corpora.
https://t.co/fCjphP59Gk
Yesterday, I presented our work on improving ASR error correction through novel data filtering, which improves robustness in various OOD settings.
Thank you to those who came by, and please reach out for any questions/comments about this work! #EMNLP
https://t.co/TkcMDKSgq0
Tomorrow 2:00-3:30pm, Aashka will be presenting our work on building efficient and effective encoder models for scientific applications (joint work with IBM and NASA). Please come join!
I’m presenting our poster on INDUS: Effective & Efficient Language Models for Scientific Applications at #EMNLP2024 tomorrow from 2:00-3:30pm at Riverside. Drop by to say hi- we have open source models & datasets!
https://t.co/QlQxei6F3K
Joint work between @IBMResearch & @NASA
ASR error correction datasets can be quite noisy, often requiring incorrect, unnecessary, or uninferable corrections. We propose two fundamental criteria to ensure data quality and show that our approach significantly reduces overcorrection while improving correction accuracy.
2 papers (1 first, 1 co-author) have been accepted to the EMNLP 2024 Industry Track 🎉
First author one was rated as Accept-top-20%. More details coming soon, see you in Florida!