I had a wonderful time presenting my work: *Finding the SWEET spot* this week at #ACL2023NLP
Thank you to everyone who took interest and stopped by🙏
Link to paper - https://t.co/gzfsI97Cvi
This work was done in collaboration with @MichaelHassid @JonathanMamou & @royschwartzNLP
Excited to present our new work on Adaptive Inference methods at #ACL2023NLP! Our paper uncovers fascinating insights about the Multi-Model and Early-Exit approaches.
Work done with @MichaelHassid, @Jonatha25240734 and @royschwartzNLP.
🧵1/9
https://t.co/y3IX4Cue3x
Read🧐, Look 👀 or Listen🎧?
What’s needed to solve a multimodal dataset.
📣Excited to share our two-step method that maps each instance in a multimodal dataset to the modalities required for processing it.
w. @YonatanBitton@royschwartzNLP
Paper📄: https://t.co/BdKj8wwgnt
1/5
We can't wait to present our work at #ACL2023NLP! For more details and additional results, check out our paper on arXiv:
https://t.co/y3IX4Cue3x
Code available here
https://t.co/xhkXnGngeo
🧵9/9
#NLP#AdaptiveInference#AIresearch#LLMs
Excited to present our new work on Adaptive Inference methods at #ACL2023NLP! Our paper uncovers fascinating insights about the Multi-Model and Early-Exit approaches.
Work done with @MichaelHassid, @Jonatha25240734 and @royschwartzNLP.
🧵1/9
https://t.co/y3IX4Cue3x
SWEET outperforms both methods in the early part of the speed-accuracy tradeoff while maintaining comparable results at slower speeds. 📈
We are excited to see our results motivate further research into fine-tuning algorithms tailored to the unique Early Exit architecture.
🧵8/9
Looking forward to seeing our findings reflected in future works.
Much more details and results in our paper, check it out:
https://t.co/uiKdi5oJou
(n/n)
How much does Attention actually attend? Apparently, not as much as you might have thought.
New paper in Findings of EMNLP with Michael Hassid, @haopeng_nlp @wittgen_ball @ivanspmontero@nlpnoah@royschwartzNLP
https://t.co/uiKdi5oJou
#EMNLP_2022
(1/n)
- PAPA also reveals a clear trend between the models’ performance and their relative reduced score, suggesting that better-performing models use their attention mechanism more.
(5/n)