New Preprint! π
w/ @EberleOliver @ashk__on and @NeelNanda5
Can we speed up finding important components in LLMs without losing faithfulness?
β Attribution patching: fast but noisy
β Integrated Gradients: more faithful but costly
β RelP: combines the best of both worlds π§΅
I am at #NeurIPS2025 in Mexico City!
Iβll present βRelP: Faithful & Efficient Circuit Discovery in Language Models via Relevance Patchingβ today at the ResponsibleFM workshop!
π Location: Don Alberto 1
β° Time: 5 p.m. β 7 p.m. CST
Donβt forget to come by our poster!
π Visit our #NeurIPS posters @NeurIPSConf!
Meet and interact with our authors at all locations β San Diego, Mexico City, and Copenhagen.
Details in the thread.
πππ
Do MLLMs really understand temporal dynamics in a video?
In our #NeuRIPS2025 paper we show that current Video-LLMs have critical limitations in temporal understanding.
To this end we proposed Stacked Temporal Attention in Vision Encoders: STAVE
https://t.co/yUvQXIrrvA π§΅π1/8
π’ Call for Papers β @LSHVU@NeurIPSConf 2025
Submissions are now open for the 7th Workshop on Holistic Video Understanding: Toward Video Foundation Models
Join us to explore multimodal reasoning, scaling laws, and evaluation for generalist video AI.
https://t.co/i2KSW82HyR
New Preprint! π
w/ @EberleOliver @ashk__on and @NeelNanda5
Can we speed up finding important components in LLMs without losing faithfulness?
β Attribution patching: fast but noisy
β Integrated Gradients: more faithful but costly
β RelP: combines the best of both worlds π§΅
@thebasepoint I appreciate you raising this. AttnLRP paper (https://t.co/ucU2vSme5w) proposes specific rules for handling attention layers. Our current RelP implementation supports only a subset of existing LRP rules, but it should be extended to include more within TransformerLens.
This was a lovely paper! Mech interp often uses attribution patching, gradient based approximations to principled casual interventions. But it's often inaccurate. Farnoush found that LRP, an XAI technique for better gradient based approximations, massively improves accuracy!
New Preprint! π
w/ @EberleOliver @ashk__on and @NeelNanda5
Can we speed up finding important components in LLMs without losing faithfulness?
β Attribution patching: fast but noisy
β Integrated Gradients: more faithful but costly
β RelP: combines the best of both worlds π§΅
β οΈ achieves comparable faithfulness to Integrated Gradients (IG) for sparse feature circuit discovery in Pythia-70M without IGβs extra computational cost.
What an unforgettable experience at #NeurIPS2024 in beautiful Vancouver! π Grateful for the chance to present my paper, attend amazing talks, and connect with brilliant minds. The ideas, energy, and city vibes made it truly special. Until next time, @NeurIPSConf !
Join us tomorrow at our poster! π
Iβll present our work, "MambaLRP: Explaining Selective State Space Sequence Models," at #NeurIPS2024.
π Thursday, December 12 π 4:30 PM PST π East Exhibit Hall A-C, Poster #2804
Looking forward to seeing you there!