Three papers from our lab at #NeurIPS2026, spanning LLM reasoning faithfulness, belief formation, and interpretability of MoEs:
⭐ Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
https://t.co/IxmJMz4Uyb
We introduce a benchmark for CoT faithfulness metrics, and find that existing metrics are either close to random or prohibitively expensive to run
@GurYoav@anmarasovic
⭐ Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models
https://t.co/S6hl0PrtgI
An interdisciplinary collaboration applying the HOT-3 indicator of consciousness to LLMs using interpretability tools, revealing surprising patterns in belief formation and action selection in LLMs
@noam_steinmetz@GoldsteinYAriel@Liad_Mudrik
⭐ Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
https://t.co/RepFOlWAne
We find a geometric coupling between routers and their corresponding experts in MoEs, explaining how routers form assignments that support an effective division of labor
@sageaarc@noyahochwald
@boardyai@megamor2 Thanks for asking! Our theoretical result doesn't depend on model size or expert count. We also saw the coupling in olmoe, qwen3 and dsv4-flash, with 64, 128 and 256 experts. Working on a revision with these results!
Excited to present today at ICML our work:
From Directions to Regions
where we show that LLM activations are better understood not as isolated directions, but as local low-rank geometries — and introduce an unsupervised method to uncover them.
📍 Thu, July 9 · Session 8 · 17:00–18:45 · Poster 3409
Also I’ll be around all week, happy to chat or grab coffee. Feel free to reach out!
I’m excited to be at #ACL2026 today presenting our paper on how matrix factorization can reveal compositional groups of neurons and hierarchies in MLPs!
📍 If you’re around, stop by our poster in Session C, 16:00–17:30. I’d love to chat!
Here we go! ✈️ 🇺🇸🇰🇷 If you're attending #ACL2026 and/or #ICML2026 check out recent work from our group with collaborators:
📍Jul 5 14:00-15:30: Matan will present his TACL work "Universal Jailbreak Suffixes Are Strong Attention Hijackers" @matanbt@mahmoods01
https://t.co/8YmWXTFA7N
📍Jul 5 16:00-17:30: My student Or will show how matrix factorization can reveal compositional neuron groups and hierarchies in MLPs. @OrShafran
https://t.co/kukRY6RsDU
📍Jul 6 9:00-10:30: Tomer will talk about "MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents" @TomerWolfson
https://t.co/MRkPTDcOJH
📍Jul 7 14:00-15:45: Hadas will present our position: "Interpretability Can Be Actionable" @OrgadHadas
https://t.co/0jy9aeruqe
(will also appear at the MI workshop)
📍Jul 8 10:30-12:15: I will be presenting a position work with @_galyo and @ymatias on metacognition as a way to move beyond hallucinations
https://t.co/tyBBL6Rlpz
📍Jul 9 17:00-18:45: Or (after a transpacific flight from ACL) will present MFA: an unsupervised approach to disentangle representations of LLMs at scale using local geometry @OrShafran
https://t.co/uSkeXBhmqe
(will also appear at the MI workshop)
New Preprint📢
Removing knowledge from LLMs is hard. Preventing models from relearning it is even harder.
In our new paper with @megamor2 and @OrShafran, we show that existing erasure methods have a blind spot: token embeddings.
The solution? EMBER🔥
🧵👇
What's in a neuron? 💫 (an atypically long, almost personal post)
Neurons in LMs have always been a fascinating object to study. I've been studying them since 2020, viewing them as key-value memory cells, analyzing what they capture in vocabulary space, and how they compose together to form features.
https://t.co/jHiZ7wLmrH
https://t.co/tYgtR6vxNi
https://t.co/eZfHTnVrgp
https://t.co/kukRY6RsDU
But many neurons still remain opaque! They do many things.
Our recent work led by @AsafAvrahamy tackles this challenge by decomposing neuron weights in vocabulary space. We do this by taking the neuron weight vector and learning different ways to rotate it (just a bit) to reveal monosemantic vocabulary channels that it captures. The nice thing about our method ROTATE is that it's data-free and super efficient, relying only on vocabulary kurtosis as a search signal.
I've been thinking about this idea since 2024, proposed it to multiple students, but only Asaf was brave enough to take this ;)
Very happy with the final outcome. Check out the paper! 👇
https://t.co/oHLMAtY1F0
Can we tell when LLMs are being unfaithful in their chains of thought?
We evaluated 8 methods claiming to do this, and found that most perform near chance!
But evaluating this requires us to have ground-truth labels for CoT faithfulness. How can we obtain these?
@muhammadshehad2@owenjonesjourno Hey Muhammad, I'm sorry for your loss. The reality is straightforward: steer clear of terrorism, and the Israeli Defense Forces will have no reason to target you.