Presenting our #ICASSP2020 paper, "EUROPARL-ST: A MULTILINGUAL CORPUS FOR SPEECH TRANSLATION OF PARLIAMENTARY DEBATES".
New resource for Speech Translation. 6 languages (En, Es, De, Fr, Pt, It) as both src and tgt.
https://t.co/0C5I7tB2Xe
https://t.co/KyxbDd8VnA
#nlp#NLProc
@bnjmn_marie Oh man, both this and your previous finding about gradient accum are shaking my foundations. Do you know the cause? There is a paper from @sarapapi about padding issues in different Conformer implementations https://t.co/4Z3jj9sEqk. Is this something similar?
@bnjmn_marie Does SmoILM have any strange architectural choices that are different from the standard Transformer? Because I agree with you, mathematically they should be the same. The other only explanation that I can come up is due to the random ordering of the samples.
Discovery of the day:
HF Transformers' gradient accumulation can significantly degrade the results compared to full batch training!
E.g.
I expected per_device_train_batch_size=1 and gradient_accumulation_steps=32 to be equivalent to per_device_train_batch_size=32 and gradient_accumulation_steps=1... but the former is much worse.
I ran several experiments with SmolLM-135M and Llama 3.2 1B. Using gradient accumulation consistently poorly performs for small batch sizes (same seed, single GPU).
Do I misunderstand something here?
I have no idea why gradient accumulation would have such a negative effect.
@bnjmn_marie Wow, that's a really huge bug if true. I will definitely be watching this closely as it has significant implications for those of us that don't have access to huge GPU clusters.
Had a lot of fun today presenting our paper and talking with people at #ACL2024NLP :)
If you missed my oral presentation at the Machine Translation session, have some lingering questions or just want to chat for a bit, I will be tomorrow at 10:30h @ In-Person Poster Session 4 !
Currently on my way to Bangkok to attend @aclmeeting#ACL2024 to represent our group @mllpresearch !!!
I will be presenting our TACL paper "Segmentation-Free Streaming Machine Translation" !
This is my first conference so I'm very excited to meet everyone in the community :)
I'm currently working as an Applied Scientist at @apptek_mclean, where I will keep working on exciting Speech and NLP projects. Looking forward to building great things together!
I defended my PhD thesis 2 weeks ago. Now that the dust has settled, I took some time to write down how everything turned out: https://t.co/HJeNhvFwkv. I also recorded a mock PhD defense in case you are interested: https://t.co/UHnAprBnLU
I really enjoyed my time at UPV, and I will always cherish the years working with my advisors and the rest of the @mllpresearch team, but it was time to move on from academia.
@bnjmn_marie Sorry to hear this. Supposedly there is a way to download and store you driving license on your phone: https://t.co/OB4zA8cTKx I have not tested this, but it might be something worth trying if what you need is just a replacement license.
What have I learned about developing ML research software during my PhD? My latest post: https://t.co/FXLRhFFCiU discusses some of the critical issues I encountered, and how can you try to overcome them.
Congratulations to my dear friend @raruidol, who recently defended his PhD thesis. Ramon produced an impressive amount of research, and I have been lucky to collaborate with him in VivesDebate-Spech(https://t.co/SAbl4puTIm). Looking forward to hear about your next achievements!
👨🏻💻 Ramon Ruiz Dolz, #VRAIN member, has succesfully defended the following thesis about "Computational Argumentation for the Automatic Analysis of Argumentative Discourse and Human Persuasion".
⬇️ Check out in our website for more information!
https://t.co/7N75e4OU0W
It is admirable to apply the precautionary principle, and build & deploy transformative AI technology with exceptional care. But how is that best achieved by distracting from very real AI risks by making a remote, fanciful risk of extinction from AI a global priority? 🤔
@osainz59 Even if they are not strictly LLMs, any pre-trained model is vulnerable to this. For example, one of my students has found strong evidence of data contamination in massively multilingual MT models.
@osainz59 Thanks for taking the time to do this analysis. This is something that I have suspected for some time, but it is nice to see some evidence. As always, more focus needs to be put into proper evaluation setups.