Is dataset debiasing the right path to robust models?
In our work, “Fighting Bias with Bias”, we argue that in order to promote model robustness, we should in fact amplify biases in training sets.
w/ @royschwartzNLP
In #ACL2023NLP Findings
Paper: https://t.co/zYP6FuVpe9
🧵👇
Can we tell when LLMs are being unfaithful in their chains of thought?
We evaluated 8 methods claiming to do this, and found that most perform near chance!
But evaluating this requires us to have ground-truth labels for CoT faithfulness. How can we obtain these?
Follow the Flow was accepted to #ACL2026 main! 🎉
T2I models often struggle with text-image alignment issues. These are usually addressed at the diffusion stage. We offer a different POV: analyzing and intervening at the textual encoding.
📄 https://t.co/KMyW4juX17
Have LLMs become supervised learners (once again)?!
In our new paper, we argue that current LLMs’ post-training methods have effectively reverted to the "pre-train then fine-tune" era, explicitly tailoring models to desired behaviors.
1/n
Fine-Tuning LLMs on New Knowledge Encourages Hallucinations. (@zorikgekhman)
But why? We found something unexpected:
1M facts about city-like names →hallucinations explode.
1M facts about random identifiers →near zero!
Same model. Same number of facts. Only the names change.🧵
Ever used a top-ranked LLM that just... felt wrong for you?
You’re not alone. Instead of leaderboards, many of us turn to "vibe-testing" - manually comparing models to our own needs. But can we turn these feelings into a structured evaluation?
New paper: "From Feelings to Metrics" 🧵
New paper: LLMs encode harmful content generation in a distinct, unified mechanism
Using weight pruning, we find that harmful generation depends on a tiny subset of the weights that are shared across harm types and separate from benign capabilities.
🧵
🚨New paper alert🚨
🧠
Instruction-tuned LLMs show amplified cognitive biases — but are these new behaviors, or pretraining ghosts resurfacing?
Excited to share our new paper, accepted to CoLM 2025🎉!
See thread below 👇
#BiasInAI#LLMs#MachineLearning#NLProc
🚨 New paper!
We present CHIMERA — a KB of 28K+ scientific idea recombinations 💡
It captures how researchers blend concepts or take inspiration across fields, enabling:
1. Meta-science
2. Training models to predict new combos
https://t.co/o2Y31xElHi
👇 Findings & data:
The longer reasoning LLM thinks - the more likely to be correct, right?
Apparently not.
Presenting our paper: “Don’t Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning”.
Link: https://t.co/Zsp3BD0TU5
1/n
Heading to @iclr_conf ✈️🧩 ‘Tokens→Words’ shows how LLMs build full‑word representations from sub‑word tokens and offers a tool for vocab expansion. 🚀
See our #ICLR2025 poster ‑ 26.4, 15:00‑17:30.
📄 https://t.co/yXvRvjjr0E
🔗 https://t.co/mTBlktKerQ
👇
[📄] Are LLMs mindless token-shifters, or do they build meaningful representations of language? We study how LLMs copy text in-context, and physically separate out two types of induction heads: token heads, which copy literal tokens, and concept heads, which copy word meanings.
✨ Ever tried generating an image from a prompt but ended up with unexpected outputs? Check out our new paper #FollowTheFlow - tackling T2I issues like bias, failed binding, and leakage from the textual encoding side! 💼🔍
https://t.co/jTNgec28hw
https://t.co/orB0Y7iW1S
🧵[1/7]
Hallucinations are a subject of much interest, but how much do we know about them? In our new paper, we found that the internals of LLMs contain far more information about truthfulness than we knew! 🧵
Project page >> https://t.co/xcQ2GNvy8H
Arxiv >> https://t.co/JCbwIexFpp
📢Paper release📢 :
🔍 Ever wondered how LLMs understand words when all they see are tokens? 🧠
Our latest study uncovers how LLMs reconstruct full words from sub-word tokens, even when misspelled or previously unseen.
https://t.co/Ur9eBn8yBO (preprint)
👀 👇
[1/7]
In which layers does information flow from previous tokens to the current token?
Presenting our new @BlackboxNLP paper: “Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers”
https://t.co/aNO7fKxXix
1/n
Ever skimmed an article, pinpointing key info, and wished for a tailor-made summary without crafting it yourself?🤔
Introducing SummHelper: your go-to for personalized summarization. 📜✏️
w/ Niv Nachum @pyshmulik@obspp18 Ido Dagan
1/n
📢 New paper alert! 📢
Thrilled to announce `Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias'.
Do instruction tuning and RLHF amplify biases in LMs? 🧵
Check it out https://t.co/tAylyaQWQi
W @boknilev@GabiStanovsky and N. Rosenfeld.
The data used to train an AI model is vital to understanding its capabilities and risks. But how can we tell whether a model W actually resulted from a dataset D?
In a new paper, we show how to verify models' training-data, incl the data of open-source LMs!https://t.co/DKQf4eoC5y