๐งตThread (1/9): People increasingly turn to AI for personal advice, and LLMs tell them what they want to hear, endorsing the user far more often than humans do.
The problem: people become less willing to repair their relationships after conflicts.
We propose PlurPO, a method to construct a pluralistic preference dataset to post-train models to mitigate this social sycophancy.
๐ Does post-training teach LLMs new agentic skills, or just sharpen what they already know?
[1/11] Introducing Sharpening Tax ๐งพ to measure the hidden cost of post-training!
Weโre excited to have a guest speaker: Dayeon (Zoey) Ki from University of Maryland to give us a talk about auditing multilingual bias across the model pipeline. It will be in ICCS X836 on Friday, August 14 at 2 PM.
We are so excited to host the 2nd edition of the Trustworthy AI for Good (AI4GOOD) Workshop at #NeurIPS2026 in Paris, France ๐ซ๐ท
We aim to bring together the AI safety, AI for social good, and AI policy/governance communities to connect what models can do as individual systems with what they do when deployed across populations.
For full details, welcome to visit our website at:
๐ https://t.co/BNWv2hNe6r
๐December 12, 2026 ยท Paris ๐ซ๐ท
We have an amazing lineup of speakers and plans. More details ๐งต
We're delighted to host the Trustworthy AI for Good (AI4GOOD) workshop at @NeurIPSConf in Paris ๐ซ๐ท
๐ข We're recruiting reviewers, plus an award for top reviewers ๐
If you're working on AI Safety, Alignment, AI for Science or related topics, we're looking for you.
Feel free to express your interest here: https://t.co/o1MvkLTjWz
We're good at measuring what alignment does to LLMs. We're much worse at describing how it happens over time. Our new paper argues alignment research should borrow predictive theories from physical sciences; using crystallization as a case study. ๐งต
๐โจ The Trustworthy AI for Good Workshop at ICML 2026 was a huge success! โจ๐
We are incredibly grateful to everyone who joined us in Seoul ๐ฐ๐ท for a full day of inspiring research, thoughtful discussions, and meaningful connections around building AI systems that are both trustworthy and socially beneficial. ๐ค๐๐ก
A heartfelt thank-you to our incredible keynote speakers:
๐ค @Yoshua_Bengio, @OanaIgnatRo, @jzl86, @maksym_andr, Jenny Ni, Naman Goyal
Your insights, perspectives, and thought-provoking ideas made the workshop truly memorable. ๐๐ง
A special thank-you to our outstanding panelists:
๐ฌ @MilindTambe_AI, @lrhammond, Gopal Sarma, @ARGleave
And to our wonderful moderator, @ZhijingJin, for leading such an engaging and insightful conversation. ๐๏ธ
We are also deeply grateful to all our oral and poster presenters, reviewers, organizers, volunteers, sponsors, and every attendee who contributed questions, ideas, feedback, and enthusiasm throughout the day. ๐
The energy of this community reinforced the importance of bringing together:
๐ก๏ธ AI safety
๐ AI for social good
โ๏ธ AI policy and governance
๐ค Cooperative AI
๐ฐ Information integrity
๐๏ธ Civic discourse
Thank you for making Trustworthy AI for Good @ #ICML2026 such a memorable and impactful event! ๐
We look forward to continuing the conversation, strengthening this community, and building on the connections formed at the workshop. ๐ค
#ICML2026 #AIForGood #TrustworthyAI #AISafety #ResponsibleAI #MachineLearning #CooperativeAI #AIResearch #TechForGood #ArtificialIntelligence
Excited for our "Trustworthy AI for Good" (AI4GOOD) Workshop at #ICML2026! As AI agents increasingly affect our lives, it is key to bridge #ResponsibleAI, social good, and governance. Letโs build solutions together!
โฐ Submission deadline: April 30, 2026 (AoE)
๐๏ธConfirmed speakers: @Yoshua_Bengio, Joel Z. Leibo (@jzl86), Maksym Andriushchenko (@maksym_andr), @OanaIgnatRo [More to come!]
๐July 10-11, 2026 ยท Seoul๐ฐ๐ท
๐ https://t.co/2NsL7jFpso
๐ Submit: https://t.co/RcOwxoPRtS
๐ฃ Be a reviewer: https://t.co/8tRAm3onbl
can synthetic training beat RAG in data-constrained domains?
we suggest a simple recipe for better synthetic training:
- Synth Mixed Training: train on both synth QAs and synth docs
- Focal Rewriting: rewrite docs with targeted topic prompts
results:
- beats RAG by +2.6% on QuaLITY
- improves to +4.4% with Focal Rewriting
- reaches +6.7% when combined with RAG
Paper: https://t.co/0rnfqUH5nE
So excited that our work is on the cover of Science!!!ย We find that AI models overly affirm users, even when they describe harmful actions. Advice from sycophantic AI made people more self-centered, yet people prefer and trust it more, which may promote this model behavior.
Thrilled to share that two of our works will be presented at EMNLP 2025!
1๏ธโฃTitle: Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
Paper: https://t.co/Z5duveIAQz
Date: Friday, Nov 7 at 14:00 - 15:30
Location: Hall C
TL;DR: We introduce MUN to test multimodal abductive reasoning in odd-to-ordinary and ordinary-to-odd scenarios, highlight gaps from human reasoning, and show that retrieval-based in-context learning yields significant gains.
โ ๏ธDifferent models. Same thoughts.โ ๏ธ
Todayโs AI models converge into an ๐๐ซ๐ญ๐ข๐๐ข๐๐ข๐๐ฅ ๐๐ข๐ฏ๐๐ฆ๐ข๐ง๐ ๐, a striking case of mode collapse that persists even across heterogeneous ensembles.
Our #neurips2025 ๐&๐ ๐๐ซ๐๐ฅ ๐ฉ๐๐ฉ๐๐ซ (โจ๐ญ๐จ๐ฉ ๐.๐๐%โจ) dives deep into this phenomenon, introducing ๐๐ง๐๐ข๐ง๐ข๐ญ๐ฒ-๐๐ก๐๐ญ, a real-world dataset of 26K real-world open-ended user queries spanning 17 open-ended categories + 31K dense human annotations (๐๐ ๐ข๐ง๐๐๐ฉ๐๐ง๐๐๐ง๐ญ ๐๐ง๐ง๐จ๐ญ๐๐ญ๐จ๏ฟฝ๏ฟฝ๏ฟฝ๐ฌ ๐ฉ๐๐ซ ๐๐ฑ๐๐ฆ๐ฉ๐ฅ๐) to push AIโs creative and discovery potential forward.
Now you can build your favorite models to be truly original, diverse, and impactful in the open-ended real world.
๐Paper: https://t.co/b4Ml1ZMJNQ
๐Data: https://t.co/YZgV6ApPjJ
We also systematically reveal Artificial Hivemind across:
๐ฅ ๐๐๐ง๐๐ซ๐๐ญ๐ข๐ฏ๐ ๐๐๐ข๐ฅ๐ข๐ญ๐ข๐๐ฌ: not only do individual LLMs repeat themselves, but different models produce strikingly similar content, even when asked fully open-ended questions.
๐ฅ ๐๐ข๐ฌ๐๐ซ๐ข๐ฆ๐ข๐ง๐๐ญ๐ข๐ฏ๐ ๐๐๐ข๐ฅ๐ข๐ญ๐ข๐๐ฌ: LLMs, LM judges, and reward models are systematically miscalibrated when rating alternative responses to open-ended queries.
(1/N)
๐คโก๏ธ๐ Post-training made LLMs better at chat and reasoningโbut worse at distributional alignment, diversity, and sometimes even steering(!)
We measure this with our new resource (Spectrum Suite) and introduce Spectrum Tuning (method) to bring them back into our models! ๐
1/๐งต
๐จ๐๐๐ work on ๐ฌ๐๐๐ฅ๐๐๐ฅ๐ ๐จ๐ฏ๐๐ซ๐ฌ๐ข๐ ๐ก๐ญ for controversial claims! ๐จ
๐๐;๐๐: AI debates help people with ๐๐ข๐๐๐๐ซ๐ข๐ง๐ ๐ฉ๐ซ๐ข๐จ๐ซ ๐๐๐ฅ๐ข๐๐๐ฌ better assess the ๐ญ๐ซ๐ฎ๐ญ๐ก in controversial casesโeven when their initial beliefs are inaccurateโshowing a promising step toward augmenting human judgment with AI in an enabling way!
Check out more details in our paper: https://t.co/ZBWqhaUpLJ
๐New Paper!
https://t.co/pyiwfwTkuL
While fact verification is essential to ensure the reliability of LLMs, detailed analysis of fact verifiers remains understudied.
We present several findings based on our revised dataset, along with practical guidance to improve the models.