@PMinervini Similarly, because of incorrect tagging of subtitle languages on youtube, if you just prompt Whisper with the expected language and it turns out to be English, it'll sometimes do a bad translation *from* English.
@PMinervini That said, the majority of "hallucinations" are merely the result of poor filtering of the input data. If you give Whisper a long audio file with a long enough piece of silence in it, you'll typically get something like "Thanks to my supporters on https://t.co/8MVPF3MJI9"
The swimming pool in a Polish border town has asked Czech visitors to stop getting naked in front of others in its locker rooms.
"Czechs have a more liberal approach to nudity" than Poles, reports a local newspaper https://t.co/SJYDBjGpoQ
📢 Our @sigdial paper, "Resolving References in Visually-Grounded Dialogue via Text Generation", is now also on @arxiv: https://t.co/uN720uK5xZ
#NLProc
🧵👇
``DDSP-based Neural Waveform Synthesis of Polyphonic Guitar Performance from String-wise MIDI Input. (arXiv:2309.07658v1 [https://t.co/mPAjnto8C8]),'' Nicolas Jonason, Xin Wang, Erica Cooper, Lauri Juvela, Bob L. T. Sturm, Junichi Yamagishi, https://t.co/Yw71nEJ80W
📢 Our @lrec2022 paper, "Collecting Visually-Grounded Dialogue with A Game Of Sorts", in which we introduce the collaborative image ranking task "A Game Of Sorts", is now also on @arxiv: https://t.co/8h5fnwB69G
#NLProc
🧵👇
OSI does not question Meta’s desire to limit the use of Llama for competitive purposes, but doing so takes the license out of the category of “Open Source.” https://t.co/ZlVASQ2K7G
SSL pre-training is expensive! How can academic groups ever hope to take part in creating these SOTA methods?
Our @ISCAInterspeech work shows that it is possible! A thread 🧵 on our efforts in reproducing speech SSL presented by @shinjiw_at_cmu at @ieeeICASSP#SASB2023:
'Níl aon gaol dhíreach idir "lá" agus "oíche" agus "te" agus "fuar", ach bíonn siad in úsáid le chéile chun cur síos a dhéanamh ar an am, agus ar an aimsir.'
Níl aithne ag ChatGPT ar "frithchiallach"
Our recent work "OverFlow: Putting flows on top of neural transducers for better TTS" is on arXiv! We improve the synthesis quality of our past work neural HMM by introducing an invertible post-net (normalizing flows) into the architecture.
https://t.co/bub0W8ZFkR
That awkward moment when you have way too many slides, blast through them, and look to see that the time up sign isn't there.... "Ummm... I think I'm done" is all I could come up with 🤦♂️