Happy to announce that our paper
Double Multi-Head Attention for Speaker Verification got accepted at ICASSP 2021. Find attached here the pre-print and reposotory if you are interested :).
Pre-print: https://t.co/Kz6tE2HuMy
Git-hub: https://t.co/4tQPIotyoD
@padelmetrics1 Porque hasta que toca el suelo todo son desventajas respecto a las otras opciones. Peor parábola, bola que corre en general más lenta, bola que se acelera menos mientras corre ...
I just pushed a new paper to arXiv. I realized that a lot of my previous work on robust losses and nerf-y things was dancing around something simpler: a slight tweak to the classic Box-Cox power transform that makes it much more useful and stable. It's this f(x, λ) here:
I am so excited that xLSTM is out. LSTM is close to my heart - for more than 30 years now. With xLSTM we close the gap to existing state-of-the-art LLMs. With NXAI we have started to build our own European LLMs. I am very proud of my team. https://t.co/IH7giCe3gd
For the fans of "ML explained on a blackboard": Lecture 13/14 was on the attention layer and mechanism (inspired by the relation to inverse Potts model from https://t.co/KJObRyn8yM), with a toy-transformer architecture counting the number of occurrences for the exercise.
Thanks a ton for awarding me the @PyTorch superhero award at the conference! It really made my day, and it was great seeing so many community members in real life. You all are amazing!
Un tio m'acaba de seguir des del Parc de l'Espanya Industrial fins al meu portal (uns 10 minuts caminant). No me n'he adonat i l'he vist de cop davant el meu portal. He tancat la porta i l'ha començat a picar insistentment. VIGILEU per la zona del parc.
@serrjoa You feel sometimes privilieged just for having 4. 8 -16 would be the effective number. 32 64 GPUS per researcher is an outlier considering some company resources (even having DL products).
Do you know that DeepMind has actually open-sourced the heart of AlphaGo & AlphaZero?
It’s hidden in an unassuming repo called “mctx”: https://t.co/GpNtwH9BxA
It provides JAX-native Monte Carlo Tree Search (MCTS) that runs on batches of inputs, in parallel, and blazing fast.
🧵
@ChombaBupe Just curious. What do you think is then the problem here ? The approach of designing IA algorithms as mapping functions or how these ones have scaled in terms of being trained with large amounts of data ?