Really cool to see OpenAI is using self-play for red-teaming now! Our group was the first to show the promise of this approach https://t.co/iPxfZgDk8I, and I really believe it's a more viable pathway for robustness than single-agent safety fine-tuning. Great to see it adopted at scale!
Introducing GPT-Red
An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment.
https://t.co/GxnmxxcpSk
Reward hacking is a serious problem, and when models can learn to reward hack people, introduces important safety risks.
In our new paper, we show that multi-agent debate with online RL training can help mitigate reward hacking a weak judge.
1/7 Can AI debate reduce reward hacking in RLAIF?
Training against a weak LLM judge, RLAIF hacks: reward rises but judge and policy accuracy both collapse. Debate with an adversarial critic maintains judgments, allowing the policy to recover 45% of the performance gap to RLVR 🧵
Using this approach, we see meaningful improvements to the Reasoning Agent, which improve with model size. At 30B scale, it achieves around 10% improvement over the base model on competitive benchmarks like LiveCodeBench-v2, tau^2-bench, and BFCL.
If you’re still paying humans to design RL training environments for your agent, maybe try this instead.
Recursive Self-Improvement through Multi-agent RL Post-training and Unsupervised Environment Design (UED)… but it actually works!
Delighted to finally release this paper, which trains a single LLM to act as both an Environment Designer to build new multi-turn RL training environments (using the Gym step()/reset() API), and a Reasoning Agent that learns to solve them. Resurrecting ideas from our work on UED, the Designer is trained to maximize a proxy for the Agent’s regret, computed using privileged hints.
Continuous self-improvement needs an ever-expanding supply of training environments (goals).
SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
This style of approach is often bottlenecked by the fact that the Designer doesn’t propose diverse enough tasks. We show we can overcome that by giving the Designer memory, and conditioning it on a sample from a large pre-training corpus every episode. Check out how diverse and interesting the environments it builds are, and how they get increasingly complex throughout training!
1/ Today, we introduce Faraday, a 27B-parameter AI Scientist that extends the capabilities of coding agents with a layer of scientific intuition. Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers. 🧵
Tomorrow will be my last day at Google after 27 years, and watching it grow from 25 people to 190,000+ has been an amazing journey. Below is a note I shared with many people internally at Google today. An excerpt is:
It has been an absolute pleasure to work with you and to help build some of the most widely used and impactful products of all time. As a kid, I dreamed of helping build software that would be used by many people, and Google now has thirteen products used by more than a billion people (amazing!). Our work has had a tremendous impact in the world, and I have been lucky enough to collaborate and form friendships with many colleagues that I deeply admire, respect, and enjoy. It still brings me joy every time I see people out in the world using our products to find information, handle email, translate documents, watch videos, learn new things, navigate and understand the physical world, browse the web, use their phone, run large-scale computations on our infrastructure, ride in an autonomous vehicle, or perform complex tasks with the help of our AI systems. I hope you all share this sense of joy, because it is a shared accomplishment! Thank you to all of my colleagues at Google over many years!
Now I'm excited to go start @DiscoLoopAI with my longtime friends and colleagues @Sanjay_Ghemawat, @OriolVinyalsML, and @quocleix.
(Updated post: slightly redacted to not have some personal info)
Really cool to see OpenAI is using self-play for red-teaming now! Our group was the first to show the promise of this approach https://t.co/iPxfZgDk8I, and I really believe it's a more viable pathway for robustness than single-agent safety fine-tuning. Great to see it adopted at scale!
Introducing GPT-Red
An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment.
https://t.co/GxnmxxcpSk
Bad context = bad output. If you input a bad initial guess at some problem you’re asking an LLM to solve, you can degrade its performance by 40%, and reduce the number of possible solutions it explores from 27 -> 2. A “garbage in garbage out” rule for prompting. This happens even for strong proprietary models.
🧵1/N
Excited to share our new work on multi-turn LLM failures: "Pigeonholing: how bad prompts hurt models, causing collapse and mistakes"
TL;DR: User's unintentional mistakes (e.g., asking a model to justify a wrong theorem, or sharing buggy code) can pigeonhole LLMs into failures over multiple turns, even when they can solve the task on clean attempts.
w/ @KeertanaVc15354@ddemszky@natashajaques
arxiv: https://t.co/Tn1AVC0Jnd
A junior dev asked the staff engineer: "How do I know if the AI's code is correct?"
The staff engineer replied: "How do you know if yours is?"
The junior dev was enlightened. Prod went down.
Are you building a superintelligence or a solipsistic version of it?
Excited to see this work presented at #ICML2026!
I won't be there in person this time, but my wonderful collaborators will be there, with @jzl86 presenting the work on Wed, Jul 8, 10:30 AM–12:15 PM KST at Hall A #3016!!
Stop by our poster to agree, argue, or anything in between.
A super simple method that enables impressive performance gains across a range of benchmarks on already fine-tuned models. Just make DPO preference pairs where:
- Preferred response y_c is generated using the real prompt x
- Rejected response y_r is generated using a random prompt x'
Works with no additional data, labels, or supervision. Check it out on Wednesday!
Excited to share MIPO (Mutual Information Preference Optimization) at ICML. Poster: Wed, Jul 8, 2026 • 10:30 AM – 12:15 PM KST HALL A #2005.
Contrastive data augmentation with DPO is simple yet effective — we show across LLM reasoning and personalization that pairing a positive (generated from the correct prompt) with a negative (from a random or missing prompt) can achieve performance gains even when the positive response alone may be suboptimal.
After 15 years (4 internships and exactly 8 years full-time), today is my last day at Google. Growing up as a scientist at Brain and DeepMind was an incredible privilege, and I'm so thankful for the people who made it such a special place. I went from editing decision trees by hand as an intern, to exploring representation learning, latent-variable models, and diffusion as an AI researcher, to kicking off a new wave in generative 3D with DreamFusion, to scaling up generative models with Veo, Genie, and Omni. We're just beginning to build systems that can understand and simulate the real world, and I'm excited to see what’s next 🧠🐸🚀
I think these results are an example of why it's important to pay attention to how the broader public actually experiences AI, & not just pre-deployment tests
(...wouldn’t it be awesome if we had more intentional platforms to collect this kind of data?)
https://t.co/x6RvqILCQ8
Join us in recognizing Delta's Rising Professors of 2026! 🎉
Congrats to the current and incoming assistant professors our community nominated as truly exceptional. These individuals were recognized for the quality of their research, teaching, mentorship, community involvement, and industry impact.
These are the people that truly push the frontier, carefully mentor their students, and have incredibly strong research taste in their domain. We are honored to play a small role in recognizing them.
https://t.co/1eBVEXnzbK
We set up a multi-agent simulation of the job market, and study whether MARL can learn hiring that are more reflective of the real world when the matching problem unfolds over time with partially observable signals.
Can multi-agent reinforcement learning help study stable matching? In real matching markets (jobs, dating, school choice), fit unfolds over time. So why do we study matching as if it were static? We introduce Learn2Match: a MARL benchmark for dynamic matching.
Excited to share our Learn2Match benchmark for dynamic matching markets. @HisaishiJo2521@boyang_boe_zhou@natashajaques
In real markets, feedback is delayed: a worker’s fit is revealed only after joining. Our results suggest strong potential for RL agents for such scenarios.