This post is long overdue, which I owe to the students from the UIUC++ SRSE program. We finally wrote a blog on how we curated a set of high-quality SRE benchmark problems by turning real-world postmortem reports into reproducible SRE problems atop #SREGym. It may give a good example of community-based data curation. Check it out,
https://t.co/vCLXrf87Zk
The fundamental challenge of AI-for-SRE benchmarks is the quality of the problems, e.g., in terms of realism and the ability of challenging frontier AI. This experience enabled us to understand many faces of the challenge and reflect on important tradeoffs between realism, reproducibility, cost, and utility. We learned many lessons, which would guide the effort on building #SREGym 2.0.
Big thanks to @SaadMRP who (unexpectedly) spent his entire summer managing the program with a heroic effort.
I list all the student contributors at the end of the post. Many of them are on Twitter: @12_AbdAllAh_12@ermiasmulu19@sharq_fr@M_ABDz_@MunimThahmid@Sai_Hari_g@ibnAmjid@SuMaya971503@TalhaAsif25@tanzimhromel@tejaspkshukla@Varunihk. They did great work to achieve high-quality problems that challenge AI frontier. If you're looking for students or employees, they are good candidates!
Well deserved, @HacksonClark and the #SREGym team! AI for SRE is in an unusual situation, where no strong benchmark is available and the technical landscape is rather opaque. The fundamental challenge is to curate high-quality problems that can reflect real-world characteristics of SRE tasks. It's great to see #SREGym continuously pushing on this direction and offering an open platform for everyone, and more importantly, the active community they have built. The #SREGym Slack has 250+ members and the project has 60+ contributors. The best part is to hear from more and more practitioners who are using it.
I'm excited to share that https://t.co/eDFaNGJ0TF has been accepted to NeurIPS 2026! 🎉
Huge thanks to my co-lead @Yiming_Su3, our co-authors @SaadMRP, @lilygn6, Yifan Tian, Hans-Arno Jacobsen, @yinfang_chen , @TianyinXu, and our 60+ contributors!
We have benchmarks for agents that write code.
We built one for what happens after you deploy it.
Can AI resolve production issues? 🧵 [1/N]
الصورة الكبيرة - الأحد 9 مساء القاهرة
كيف تحصل على القبول في MIT، أحد أرقى جامعات العالم. وكيف تحصل على منحة دراسية. حوار مع د. أحمد غنيم، رئيس قسم الهندسة الميكانيكية بالـ MIT يجيب على أسئلتكم.
https://t.co/cumZP5rULe
Tons of great perspectives from a systems professor 🥹
lots of great principles in systems research 💫 can be borrowed from and apply to ai agent systems - fault tolerance, recovery, formal methods, security etc
徐天音:系统,UIUC,教授,最佳论文,Agent Infra,云计算,形式化验证,容错,纯粹,好导师 https://t.co/gtz7g9KhvB via @YouTube
If u guys are into inference engineering, do check out this.
They have got a proper roadmap and hands on exercise to build you into an inference expert.
https://t.co/Qs6ZwJL03a
A higher KV cache hit rate doesn’t always mean lower TTFT and TPOT, or higher throughput.
We studied LLM routers with different request scoring objectives under agentic and RL workloads. Here’s what we learned. 🧵
Non-PhDs: “When I retire, I’ll finally do a PhD. Research sounds so fun.”
PhDs: “If I hadn’t done a PhD, I could’ve joined one of those fancy AI labs before ChatGPT.”
Everyone romanticizes the path they didn’t take.