Excited to present our ICML 2026 poster!
Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov Games
KL regularization has become a fundamental tool in modern RL and alignment, but its statistical benefits have remained largely understudied in many settings. In this work, we provide the first logarithmic regret guarantees for Matrix and Markov Games under KL regularization.
The main challenge is that KL-regularized Nash equilibria do not admit closed-form expressions, unlike KL-regularized optimal policies in single-agent RL, making the standard logarithmic regret analysis inapplicable. To overcome this, we develop:
✅ Best-response sampling around estimated Nash policies
✅ Optimistic bonuses for Matrix Games (OMG)
✅ Super-optimistic bonuses for Markov Games (SOMG), which preserve optimism throughout backward induction under best response sampling.
More details :📄 Paper: https://t.co/DtNhLxePSK
If you're attending ICML, come by and say hi! We'd be happy to discuss the paper, answer questions, or chat about RL, LLMs, game theory, alignment, or ML in general.
📍 ICML 2026 Poster #119 (Hall A)
🗓️ July 9, 10:30 AM–12:15 PM (KST)
What do we control when fine-tuning diffusion models?
Excited to share our ICML 2026 paper "Diffusion Controller (DiffCon)": we view the diffusion reverse process as state-only control in a generalized linearly-solvable (LS) MDP. The LS-MDP view gives rises to unifying and practical RL finetuning algorithms, and a lightweight score correction suitable for both gray-box and white-box finetuning.
Excited to share our recent work! We provide a mechanistic understanding of long CoT reasoning in state-tracking: when do transformers length-generalize strongly, when they stall, and how recursive self-training pushes the boundary. 🧵(1/8)
🚨 🔥 Multi-step reasoning is key to solving complex problems — and Transformers with Chain-of-Thought can do it surprisingly well.
🤔 But how does CoT function as a learned scratchpad that lets even shallow Transformers run sequential algorithms that would otherwise require deeper architectures?
In our new work https://t.co/tSoFRsgO39, we prove that even a 1-layer multi-head Transformer trained via gradient descent can perform multi-step symbolic reasoning on trees.
What we find:
🔄 The model learns to “turn around”: first generates goal→root path, then reverses to root→goal path—all within a single autoregressive pass.
🧩 Two attention heads learn to specialize and collaborate: 🧭 one head follows the path; 🚦 one head tracks the phase and triggers the flip at the root.
@yuejiec@yuhuang42
🚨Ultimate Jailbreaking Championship 2024 🚨
Hackers vs. AI in the arena. Let the battle begin!
🏆 $40,000 in Bounties
🗓️ Sept 7, 2024 @ 10AM PDT
🔗Register Now: https://t.co/F0w51xMD8U
No LLM is secure! A year ago, we unveiled the first of many automated jailbreak capable of cracking all major LLMs. 🚨
But there is hope?!
We introduce Short Circuiting: the first alignment technique that is adversarially robust. 🧵
📄 Paper: https://t.co/hY7koqrLyl