Do you care about metrics for model explanations?
How can human explanations guide your model?
Are explanations *sufficiently* informative?
Follow my oral pres. at IJCNLP-AACL 2025 tomorrow, or check out our paper if you're interested!
https://t.co/aCwWCjglzi
⚠️ Citations from prompting or NLI seem plausible, but may not faithfully reflect LLM reasoning.
🏝️ MIRAGE detects context dependence in generations via model internals, producing granular and faithful RAG citations.
🚀 Demo: https://t.co/OMeM32TNoY
Fun collab w/ @Jirui_Qi, @AriannaBisazza & @raquel_dmg! Check it out ⬇️
Also at @LrecColing and interested in #interpretability in #nlp / #xai? I'm happy to tell you something about post-hoc attribution methods and their linguistic preferences 🔻
Poster session >> from 17:30 onwards
Proceedings: https://t.co/T7bTmWZ0Rw
Did you know that token-level differences between feature attribution methods are smoothed out if we compare them on the syntactic span level?
I'll be at @LrecColing (Turin) next month to present our paper on the linguistic preferences of such methods :) https://t.co/hOUeNCHfQs
Interested in post-hoc explanation methods? Swing by our #EMNLP2023 poster today at 10:30 to chat about their (dis)agreement and the implications of selecting the top-k most important tokens!
(https://t.co/M8jywnyihQ)
Just wrapped up our presentation at #MMNLG! Big thanks to everyone for the thought-provoking questions, and kudos to @bertugatt for the seamless organisation. Can't wait to dive deeper into this work!
#SIGDIALxINLG2023
Shout-out to @hmohebbi75 et al. from our #InDeep consortium for their awesome work "Quantifying Context Mixing in Transformers", introducing Value Zeroing as a new promising post-hoc interpretability approach for NLP! 🎉 Paper: https://t.co/ebwuty4CLL
Some smooth work by @gsarti_ and colleagues from the #InDeep project in releasing a tool for interpreting sequence generation models. Keep up the development! @InseqDev
After a year of restless development, I'm finally happy to announce Inseq, a new tool to democratize post-hoc interpretability of sequence generation models 🐛 https://t.co/YlxGuW4bVO #nlproc#xai
Some highlights 👇 1/
@annargrs@emnlpmeeting Virtual attendance to workshops was fine. As for gathertown, I would have liked to see the names of the papers instead of 'paper #' and perhaps more authors present at their posters during the dedicated time slot.
So far today: I've been able to virtually attend co-occurring @NllpWorkshop and #BlackboxNLP and listen to some inspiring talks. 🧙 Ito hard skills, I learned how to mute/unmute individual zoom chrome tabs
Moonlit (https://t.co/Vde4m741HV) just released the interview they had with me! Check it out if you want to get some idea of what my research on #interpretability and (legal) #argmining is about! https://t.co/46Vv6YH0id
What can we learn about how named entities are represented in language models by substituting them for different ones and observing how predictions change?
Come see my poster this Thursday December 8th @ BlackboxNLP 2022 in Abu Dhabi (and online).