Sycophancy is the tendency to produce responses that appeal to users at the expense of good judgment.
Ironically, the very stage meant to align models with human intent can make them less trustworthy as advisors
Pluralistic alignment is thriving as a research agenda yet failing at its goal: making the AI systems people actually use more pluralistic🌈
🚨New position paper: we argue adoption in deployed models should be the fields main goal & we provide a roadmap of how to get there
🧵1/
📢Call for abstract submissions: 2026 Workshop on Human-AI Complementarity for Decision Making at CMU
This year's theme: Dynamic Human-AI Alignment. How do we design AI systems that don't just emulate static human preferences, but actively and effectively coordinate with humans over time?
📅Workshop Dates: September 24-25, 2026
🗺️Workshop Location: Pittsburgh, PA, USA
⌛️ Application Deadline: Friday, July 17, 2026
Funding for travel/lodging available to support speakers + student presenters.
I'm in Seoul attending #ICML2026 this week, presenting work on LLM sycophancy and evaluation!
On Tuesday morning (#3309) I'll be presenting our work on how RLHF fine-tuning can make models more sycophantic (work with @ArielProcaccia and Gerdus Benadè)
When agreement with the user's framing is overrepresented among high-reward outputs relative to responses that correct the user's premise, optimization makes the model more agreeable.
A (particular) annotator bias is enough for a Bradley-Terry reward to satisfy this condition
https://t.co/3lZzTS4AWN
Excited to attend #ICLR25 this week. My DMs are open, feel free to drop a message to talk about anything related to optimization of deep networks.
Presenting multiple works related to second order optimization, critical batch size and diagonal preconditioning. Details below.
When I realized this today, I dug a little into the literature and thought "gah I'm scooped!".... but this recent paper actually seems to be saying something quite different:
https://t.co/8NZGGsr1jJ
(7/8)
1/n A technical thread on our results in https://t.co/Idqc1JPpIy on connecting the Shampoo optimizer and Optimal Kronecker product approximation of the the Adagrad (or Hessian) preconditioner.
Why does Shampoo work well? Our new work sheds light on this, highlighting a wide misconception about the optimizer. We show the *square* of Shampoo's preconditioner is provably near to the optimal Kronecker approximation of the (Adagrad) Hessian. See:
https://t.co/tnHNQxH7Fz
Excited to share our paper on generative social choice, which "fuses the rigor of social choice theory with the flexibility and power of generative AI." We plan to put this approach into action soon as part of @OpenAI's "Democratic Inputs to AI" program.
https://t.co/VcWiR9iTA2
We've just launched fine-tuning for GPT-3.5 Turbo! Fine-tuning lets you train the model on your company's data and run it at scale. Early tests have shown that fine-tuned GPT-3.5 Turbo can match or exceed GPT-4 on narrow tasks: https://t.co/VaageW9Kaw
@TShadmy מודלים כמו GPT4 עונים על רוב הקריטריונים בהגדרות מוסכמות של ״אינטיליגציה״ (כמו הגדרות שהתקבעו ממחקרים על בעל חיים) כמו להסיק, לפתור בעיות, לחשוב אבסטרקטי, להבין רעיונות מורכבים וכו׳
@TShadmy 1. חלק גדול מהחוקרים בתחום הזה יעידו שהיו מופתעים מאד מהיכולות של LLM.
2. זה יכול להיות מטעה לחשוב על GPT כאל מכונה סטטיסטית כי הוא לא מציית לחוקים סטטיטים כמו שהיינו מצפים.
3. ככל הנראה צריך להיות מאד ״אינטליגנטי״ כדי לחזות את המילה הבאה ברצף של טקסט רנדומלי מהאינטרנט
@liron Funny you mention that, because for years there has been an ongoing debate on the hype vs. real-world impact of deep learning, with skeptics questioning the significance of its practical and scientific applications (except for AlphaFold).
@karpathy Considering the training data's 80% '1's and 20% '0's distribution, it's reasonable to expect a 73% transition probability from '000' to '001', instead of a random 50% guess