Understanding Matching Mechanisms in Cross-Encoders
@MathiasVast1 et al. investigate how neural ranking models construct relevance signals by analyzing attention processes and extracting causal insights in cross-encoders.
📝 https://t.co/kQ6Bu8KZlz
👨🏽💻https://t.co/w9R49O1EGO
Which Neurons Matter in IR? Applying Integrated Gradients-based Methods to Understand Cross-Encoders
Adapts Integrated Gradient methods to interpret MonoBERT for IR, identifying specific neurons responsible for relevance assessment and OOD handling.
📝https://t.co/zX7t0GkAuI
Just came back from Washington and the @SIGIRConf + @acm_ictir conferences where I have had the chance to present our last paper "Which Neurons Matter in IR? Applying Integrated Gradients-based Methods to Understand Cross-Encoders" during the ICTIR day. (1/4)
Amongst other things, we observed that relevant passages and non-relevant passages are treated by different neurons, OOD predictions involve additional neurons compared to ID and the important neurons are localized mostly in 2 regions of the model (middle and last layers). (3/4)
The last Sentence Transformers release introduced GISTEmbedLoss by @avsolatorio, which allows for training models that outperform those trained by the wildly popular in-batch negatives loss (MultipleNegativesRankingLoss). Learn about it in this 🧵(links at the bottom):
All done! It's a wrap from #ecir2024. We've had a brilliant 5 days with loads of great talks, discussions, and social events.
Thanks to everyone that made #ecir2024 such a great success! We hope you all enjoyed Glasgow. Looking forward to seeing you again in Lucca for #ecir2025
If you still have the energy after all the nice presentations from today's sessions and keynote, don't hesitate to come and discuss with @MathiasVast1 during the poster session later today @ecir2024 !
@IntranetFocus@sinequa@arxiv Samples from Wikipedia also have the advantage to be freely accessible (to us for training and testing and to the research community for reproducing). It is thus more relevant, in the context of research, to experiment on these datasets.
@IntranetFocus@sinequa@arxiv Good point. Indeed Enterprise Search has its own specificities that might not be well-represented in a corpus coming from Wikipedia. Nevertheless, some features are shared which allows us to draw conclusions, even if the numbers don't always translate.
The year 2023 is drawing to a close.
Before we head off on vacation, let's take a look back at the year:
(1/5) Our ML/DL team has grown with the arrival of:
- François Yvon @yvofr, CNRS research director
- Caio Corro @UndefBehavior & Alasdair Newson, lecturers @SorbonneParis1
I am thrilled to share that my first paper "Simple Model Adaptation for Sparse Retrievers" in collaboration with @zongyuxuan3@bpiwowar and @LaureSoulier has been accepted in the "short papers track"! Special thanks to @basilevancooten too for the supervision #ecir2024