🤔 Our latest ICLR-oral, T3, studied how LLM agents get belief-trapped in long multi-turn interactions: they keep acting, but their internal task understanding drifts away. But this raises a deeper question:
Why does outcome-based RL sometimes fail to fix this — and even reinforce low-information behavior?
🔥 Our ICML 2026 paper studies this unique failure mode in agentic RL: Information Self-Locking (SeL).
🔗 Paper: https://t.co/hbY1CvaQsr
💻 Code: https://t.co/J4ipDpnUMw (merging with T3!)
🍲Quick video: https://t.co/ukJ5owcNQC
Huge thanks to all collaborators from CUHK, UCSD, ByteDance, and Georgia Tech!
Proud to see this great work from my friend! 😋
Multi-turn reasoning ≠ progress.
Nice work on identifying and fixing belief-trapped trajectories in LLM agents (T3). Congrats on the ICLR Oral 🔥
🤔Do you also often find that agents get lost in multi-turn interactions — whether in conversations, coding, or even long-form writing?
In many agentic applications, agents perform active reasoning: they interact with the environment to gather information and progressively resolve the task.
🔥 Our ICLR 2026 Oral paper studies this common failure mode in RL for LLM agents: in multi-turn active reasoning, the agent may keep interacting while making little real progress.
We characterize this rigorously and propose a simple yet effective mechanism: T3 (Truncating Belief-Trapped Trajectories).
Thank all collaborators from CUHK, ByteDance and Gatech!
🔗 Paper: https://t.co/3yhJz9Atnj
💻 Code: https://t.co/J4ipDpnUMw
Excited to attend NeurIPS 2025 with our OPTML group!
Grateful for the opportunity to learn, present, and connect with researchers working on trustworthy AI.
See you in San Diego! 🌟✈️
✈️✈️✈️Heading to San Diego for #NeurIPS2025! Thrilled to share @OptML_MSU’s “Menu of Innovations”, a showcase of our students’ great work in LLM interpretability, unlearning, reasoning safety, and model honesty (one Spotlight, two Posters, one Workshop paper, and one Rising Star Award).
Excited to meet new and old friends in San Diego! And OPTML is hiring PhD students; If you’re interested in trustworthy and scalable AI, feel free to ping me and meet up. 🚀 @zyh2022@ChongyuFan@wcsa23187
🎯 Our EMNLP 2025 Main paper
“Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills” goes live soon!
Catch us on Wednesday in Suzhou at #EMNLP2025 🇨🇳
🔗 Paper Link: https://t.co/6zm3EYbfHk
🏡 Project Page Link: https://t.co/1bBgTJ8qRW
🗓 November 5, 11:00–12:30 CST (UTC+8)
📍 Hall C, Section 2, 500-Main
🧍 I won’t be there in person — but feel free to chat with my co-authors!
🧠 The Problem
You’ve erased sensitive answers from your LRM.
But the reasoning traces, the step-by-step “thoughts” that led there, still remain.
Even after unlearning, the model can reconstruct or re-infer forgotten answers through these traces.
So the question is:
👉 Can we truly forget reasoning traces, while preserving the model’s reasoning ability?
🎯 Our Solution: R²MU (Reasoning-aware Representation Misdirection for Unlearning)
We go beyond answer-level forgetting and target the reasoning process itself.
R²MU suppresses sensitive reasoning traces while maintaining general reasoning competence.
Through representation misdirection, the model unthinks unsafe reasoning paths, while CoT supervision preserves valid reasoning skills.
⚙️ How it Works
🔄 Unthinking Loss: misaligns hidden representations of sensitive reasoning traces with randomized features.
💡 Reasoning Preservation: uses CoT datasets (like LIMO) to retain problem-solving ability.
✅ R²MU erases reasoning traces — not just answers.
✅ Preserves general reasoning and utility across diverse benchmarks.
✅ Achieves the lowest reasoning-trace leakage (RT-UA ↓) on unlearning benchmark WMDP and LRM safety benchmark STAR-1, while maintaining top reasoning accuracy on AIME, MATH-500, and GPQA.
👥 With amazing collaborators from MSU: @ChongyuFan ,@zyh2022 , @jia_jinghan , and my advisor @sijialiu17 .
🙏 Grateful to our IBM collaborators from @MITIBMLab : @NathalieBaraca1 , Dennis Wei, @p_ram_p.
Announcing 🔭✨Hubble, a suite of open-source LLMs to advance the study of memorization!
Pretrained models up to 8B params, with controlled insertion of texts (e.g., book passages, biographies, test sets, and more!) designed to emulate key memorization risks 🧵
new research from Meta FAIR: Code World Model (CWM), a 32B research model
we encourage the research community to research this open-weight model!
pass@1 evals, for the curious:
65.8 % on SWE-bench Verified
68.6 % on LiveCodeBench
96.6 % on Math-500
76.0 % on AIME 2024
🧵
Computer Use: Modern Moravec's Paradox
A new blog post arguing why computer-use agents may be the biggest opportunity and challenge for AGI.
https://t.co/vq7s73OYUg
Table of Contents
> Moravec’s Paradox
> Moravec's Paradox in 2025
> Computer use may be the biggest opportunity for AGI
> Chatbots → agents
> Internet-scale learning of human cognition
> Bits > atoms
> Enormous economic value
> Why is computer use hard for AI?
> Computer use ≠ clicks + typing
> Idiosyncratic environments
> Contextual understanding
> Tacit knowledge
> Is RL the panacea?
> Looking forward
If you are also excited about CUAs and want to do some serious work, let's chat!
New Scale research: Can smaller models reliably oversee stronger LLM agents?
We red team monitoring systems to detect covert sabotage, like agents secretly downloading sensitive information.
What if you could not only watch a generated video, but explore it too? 🌐
Genie 3 is our groundbreaking world model that creates interactive, playable environments from a single text prompt.
From photorealistic landscapes to fantasy realms, the possibilities are endless. 🧵
Thank you @INNSociety for this great honor. I am deeply grateful to my nominator, students, and collaborators who made this recognition possible. Excited to keep advancing the frontiers of scalable and trustworthy AI! @OptML_MSU
🤔Come to check Chongyu @ChongyuFan and Jinghan Jia @jia_jinghan ‘s work!
👨💻Jinghan has been an incredible mentor to work with—smart, supportive, and inspiring. He’s on the job market now, feel free to reach out and chat with him!
Excited to share our ICML’25 work on robust LLM unlearning!
Poster is on Wed, July 16 @ 4:30pm PT (E-2803) , feel free to stop by and chat with the team!
🎉 Excited to share that our OPTML lab has two papers on unlearning robustness at #ICML2025!
🧠 Plus, we’ll present our work on reasoning unlearning at the MUGen Workshop.
Come check out our posters and chat with us!
🚨 Excited to attend #ICML2025 and share our latest work (@OptML_MSU) on LLM unlearning -- think of it as AI surgery: removing harmful knowledge while preserving general utility. Catch us at:
🔹 [Paper 1] Tues, July 15 @ 4:30pm PT | E-1108
📄 Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning [🔗 https://t.co/fb7sgRiTBc]
-- Even unrelated fine-tuning (e.g., math) can achieve invariance in unlearning via disentangled task vectors, which improves unlearning robustness against general post-unlearning fine-tuning operations.
🔹 [Paper 2] Wed, July 16 @ 4:30pm PT | E-2803
📄 Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond [🔗 https://t.co/ERl2v8JsTQ]
-- Connecting adversarial unlearning to sharpness-aware optimization and general smoothness optimization
Also at #MUGen Workshop @ ICML:
🎤 Invited talk: Progress, Pitfalls & Prospects of LLM Unlearning
🧠 Oral: "Unlearning Isn’t Invisible: Detecting Unlearning Traces in LLMs from Model Outputs" [🔗 Long version: https://t.co/ZJwWX8Jc8v]
📌 Poster: "Reasoning Model Unlearning" [🔗 Long version: https://t.co/WJOcqb3O4r]
Grateful to my outstanding students (@ChongyuFan@wcsa23187@yiwei_chen_@Hi_Soumyadeep@zyh2022@jia_jinghan) and wonderful collaborators from @MITIBMLab IBM (@NathalieBaraca1 Dennis Wei @p_ram_p) and Amazon (@Mingyi552237 @anil_k_ram) for their dedication and contributions to these works!
Looking forward to reconnecting with friends and meeting new ones--come say hi and chat about building safe, efficient, and trustworthy generative models! 🤝