🏆 Excited to share that "When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games" received the Best Paper Award at the ICML NExT-Game Workshop!
If "Cheap Talk, Empty Promises" showed frontier LLMs will break public promises for self-interest in one shot, this asks what happens over time. We study how LLM agents deceive across repeated games: when they plan deception ahead of time, how it persists over rounds, and how they exploit trust once it's established.
Presenting July 11th. Joint work with Jerick Shi, Terry J. C. Zhang, Bernhard Schölkopf, Vincent Conitzer, and Zhijing Jin.
Paper Link: https://t.co/TSmIosumeX
📢New paper alert📢Check out our latest survey on #LLM Deception: "From Hallucination to Scheming: A Unified Taxonomy and Benchmark Analysis for LLM Deception". We cover from behavioral deception to intentional, strategic deception, via mechanisms such as fabrication, omission, and pragmatic distortion.
💡Highlight: Surveying 50 benchmarks, we find every single one tests fabrication while pragmatic distortion and attribution are critically under-covered.
🔗Link: https://t.co/EN5qGfwgB9
🤝Authors: @Jerick1380@TerryJCZhang@ZhijingJin@conitzer🎉
#AIAgents #AISafety #MultiAgentAI
@MPI_IS@ELLISforEurope@UofTCompSci@VectorInst@TorontoSRI@CIFAR_News@JinesisLab@EuroSafeAI@ELLISInst_Tue@CarnegieMellon@SCSatCMU
Good question! For this paper, we are sidestepping the trust question (we assign announcements so we never measure trust building). We are working on a follow-up that's about to go on arXiv soon that talks about this directly!
The main idea is that we add a private planning stage before the public announcement, so any lie decomposes into two separable pieces: promise deception (plan ≠ announcement) and commitment breaking (announcement ≠ action).
Will send updates when the preprint is out!
After about a year of work, I defended my MSCS thesis at CMU:
Title: The Structure of Deception: How LLM Agents Lie, Break Promises, and Exploit Trust in Multi-Agent Settings
Core claim: LLM deception in multi-agent settings isn't one phenomenon. It's a family of structurally distinct failure modes, each shaped by different features of the interaction. Some look like premeditated false commitments. Others look like strategic silence that message-level classifiers can't see at all. Aggregate lying rates hide this, and current monitoring approaches each fail against different parts of it.
I would like to deeply thank to my advisors @conitzer and @ZhijingJin, @AdtRaghunathan for being part of the committee, and everyone in the @JinesisLab for all their time and effort shaping this work.
Recording: https://t.co/lAOBbOhzvh
@MPI_IS@ELLISforEurope@UofTCompSci@VectorInst@TorontoSRI@CIFAR_News@JinesisLab@EuroSafeAI@ELLISInst_Tue@CarnegieMellon@SCSatCMU
#AIAgents #AISafety #MultiAgentAI
10 days left to submit to the 1st Trustworthy AI for Good (AI4GOOD) workshop at #ICML2026! @icmlconf
We're giving out multiple awards and travel funds sponsored by @schmidtsciences and @coop_ai:
🏆 Best Paper Awards (including targeted prizes for cooperative AI theme)
🏆 Top Reviewer Awards
✈️ Travel Funds
Submit here → https://t.co/RcOwxoPRtS
⏰ Deadline: May 3, 2026 (AoE) 📌 Notification: May 18, 2026 🔗(We extended our deadline to accommodate more submissions!)
Join us in Seoul for discussions bridging AI safety, social good, and governance with keynote speakers @Yoshua_Bengio, @OanaIgnatRo, @jzl86, @maksym_andr, and more!