Very glad and grateful to share this fun work, especially because of the opportunity to work with this immensely talented team--@jkminder, @niklas_stoehr, @giomonea, @wendlerch, @cervisiarius, & Ryan Cotterell 🥳. Check it out!
Can we understand and control how language models balance context and prior knowledge? Our latest paper shows it’s all about a 1D knob! 🎛️
https://t.co/698wbJ0IDZ
Co-led with @kevdududu, as well as @niklas_stoehr, @giomonea, @wendlerch, @cervisiarius & Ryan Cotterell.
Does the *way* we say something to an LLM affect whether the LLM agrees with it?
Our ACL 2026 paper - “It’s Not What You Say, It’s How You Say It”, answers with a resounding yes! ✅
https://t.co/Gdr6zrgOKL
Co-led with @ClaraKuempel, as well as @MichelleWastl and @a_stadt 🧵
The main takeaway: It's not (just) what you say, it's how you say it!
The next time you evaluate LLMs for safety or instruction-related tasks, consider varying the way you express a belief to a LM because it might just behave very differently depending on how you say it.
LLMs can accept the same false claim differently depending on its tone, certainty, and grammatical form.
Small wording changes can make LLMs accept false claims, while larger and instruction-tuned models resist them more.
Models must decide whether to trust a user’s new claim or rely on facts stored during training.
EoBench tests this choice with about 66K false claims written in 19 styles across form, evidence, certainty, and tone.
The team evaluated 18 Gemma, Llama, and Qwen models, then kept cases where each model already knew the correct fact.
Commands, child-directed wording, formal language, and authority claims persuaded models most, while weak claims and counterfactuals persuaded them least.
Across Llama and Gemma, larger models followed false context less often, and instruction tuning usually reduced that behavior.
The finding shows that prompt wording can quietly change model answers, so evaluations and product safeguards must test linguistic framing directly.
---
– arxiv. org/abs/2607.18232
Title: "It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief"
Can we understand and control how language models balance context and prior knowledge? Our latest paper shows it’s all about a 1D knob! 🎛️
https://t.co/698wbJ0IDZ
Co-led with @kevdududu, as well as @niklas_stoehr, @giomonea, @wendlerch, @cervisiarius & Ryan Cotterell.
Our new mechanistic interpretability work "Activation Scaling for Steering and Interpreting Language Models" was accepted into Findings of EMNLP 2024! 🔴🔵
📄https://t.co/VQ02Yz3Bq8
@kevdududu, @vesteinns, @cervisiarius, Ryan Cotterell and @AaronSchein
thread 👇
@yc6n13@vesteinns@niklas_stoehr@JenniferCWhite@AaronSchein @ryandcotterell by preference fine-tuning, do you mean using them as the reward function? Yes! we are actively looking into methods for controlling LMs to be more or less context-driven and this is def a direction we're investigating :)
How much does an LM depend on information provided in-context vs its prior knowledge?
Check out how @vesteinns, @niklas_stoehr, @JenniferCWhite, @AaronSchein, @ryandcotterell + I answer this by measuring a *context's persuasiveness* and an *entity's susceptibility*🧵
To answer this, we introduce measures based on mutual information for the *persuasiveness* of a context and the *susceptibility* of an entity. Intuitively, a context's *persuasion score* is how much a model's answer distribution to a query changes when provided the context. (2/n)
finally, we showcase how these scores can be useful to model practitioners in two downstream applications–friend-enemy stance detection and analyzing gender biases in models. e.g., we find for those 2 questions, enemy duos are less susceptible than friend duos! (8/8)