I'm excited to share that https://t.co/eDFaNGJ0TF has been accepted to NeurIPS 2026! 🎉
Huge thanks to my co-lead @Yiming_Su3, our co-authors @SaadMRP, @lilygn6, Yifan Tian, Hans-Arno Jacobsen, @yinfang_chen , @TianyinXu, and our 60+ contributors!
We have benchmarks for agents that write code.
We built one for what happens after you deploy it.
Can AI resolve production issues? 🧵 [1/N]
This post is long overdue, which I owe to the students from the UIUC++ SRSE program. We finally wrote a blog on how we curated a set of high-quality SRE benchmark problems by turning real-world postmortem reports into reproducible SRE problems atop #SREGym. It may give a good example of community-based data curation. Check it out,
https://t.co/vCLXrf87Zk
The fundamental challenge of AI-for-SRE benchmarks is the quality of the problems, e.g., in terms of realism and the ability of challenging frontier AI. This experience enabled us to understand many faces of the challenge and reflect on important tradeoffs between realism, reproducibility, cost, and utility. We learned many lessons, which would guide the effort on building #SREGym 2.0.
Big thanks to @SaadMRP who (unexpectedly) spent his entire summer managing the program with a heroic effort.
I list all the student contributors at the end of the post. Many of them are on Twitter: @12_AbdAllAh_12@ermiasmulu19@sharq_fr@M_ABDz_@MunimThahmid@Sai_Hari_g@ibnAmjid@SuMaya971503@TalhaAsif25@tanzimhromel@tejaspkshukla@Varunihk. They did great work to achieve high-quality problems that challenge AI frontier. If you're looking for students or employees, they are good candidates!
I'm excited to share that https://t.co/eDFaNGJ0TF has been accepted to NeurIPS 2026! 🎉
Huge thanks to my co-lead @Yiming_Su3, our co-authors @SaadMRP, @lilygn6, Yifan Tian, Hans-Arno Jacobsen, @yinfang_chen , @TianyinXu, and our 60+ contributors!
We have benchmarks for agents that write code.
We built one for what happens after you deploy it.
Can AI resolve production issues? 🧵 [1/N]
@melissapan@Yiming_Su3@SaadMRP@lilygn6@yinfang_chen It's one of my favorite impacts of our work! We have many enthusiastic users who actively give back to our work, and it's one of the most rewarding parts of the project.
@Yiming_Su3@SaadMRP@lilygn6@yinfang_chen Come find out if AI resolve production issues.
https://t.co/HdYZGsunKR
arXiv: https://t.co/69lrfINsw9
GitHub: https://t.co/xNS7dNcxQ1 Leave a ⭐️!
See you at NeurIPS in Sydney! 🇦🇺
9/ SREGym is completely open source, and the project has grown far beyond what we originally described in the paper thanks to an amazing community of contributors.
Our goal is simple: build THE benchmark for AI agents operating software in production. We're just getting started!
8/ And agents have gotten a lot better since we ran those experiments.
That's exciting—but it also means the benchmark has to get harder.
We're working on the next generation of SREGym: harder incidents, stronger evaluation, and more realistic production environments.
More on that soon 👀
7/ We also found that SOTA agents fail in some surprising ways.
Agents often take a greedy approach: they see an anomaly, form a hypothesis, and immediately try to fix it.
Even when later presented with evidence of the actual fault, they can remain anchored to their initial hypothesis.
There's still a lot of work to do before we can reliably deploy these agents in production.
6/ In the paper, we evaluated frontier agents across 90 production failures.
One of the clearest findings: performance depends heavily on *what kind* of production failure the agent encounters.
We observed differences of up to 40 percentage points in end-to-end resolution across failure categories.
Production reliability isn't one skill.
5/ Our problems come directly from real production incidents.
We've reproduced failures inspired by Cloudflare exhausting CPUs with a WAF rule, GKE clusters running out of IP addresses, Linux conntrack exhaustion, Kafka poison pills, admission webhook failures, and Reddit's Pi-Day outage.
Real postmortems help ground SREGym in the failures that production engineers actually encounter.
4/ And production failures can get weird.
Since the 90-problem benchmark evaluated in our paper, SREGym has grown to 125 problems—35 new problems spanning applications, Kubernetes, networking, operating systems, hardware, and more difficult issues like concurrent and metastable failures.
And we're still growing!
3/ That's SREGym.
We put AI agents inside live cloud environments, break something, and ask them to fix it.
Agents investigate the system using real observability and operational tools, diagnose the failure, and take actions to restore the application.
The task isn't "write a patch."
It's: production is broken. Fix it.
2/ Coding benchmarks have driven incredible progress in AI agents.
But writing code is only part of the software lifecycle.
As agents get better at writing and deploying software, operations become the next bottleneck.
Someone—or something—still has to keep that software running: investigating outages, navigating complex distributed systems, finding root causes, and safely restoring service.