I’ll be presenting FrontierCS: Evolving Challenges for Evolving Intelligence at #ICML 2026 in person!
FrontierCS is a benchmark for open-ended computer science problems, where agents need to improve executable solutions under objective, fine-grained evaluation rather than solve tasks with fixed known answers.
We’ll also be sharing several new research directions building on FrontierCS, including:
(1) How can we automatically synthesize open-ended problems?
(2) How can open-ended problems benchmark the scaling capabilities of agents and auto-research systems?
(3) What is the roadmap toward FrontierCS 2.0?
Poster: #61419
Come find me in person at ICML — happy to chat about long-horizon agents, auto-research, and open-ended evaluation!
We taught a brand-new mini-series this year at @SCSatCMU on Modern GPU Programming for ML Systems, as part of the ML Systems course, touching on fun questions like what data layout swizzling is, how to use 3D TMA, and state-of-the-art Blackwell programming. We released a curated online book based on the materials: https://t.co/5ZJg2lySNO check it out
We release TIRx today, a minimal compiler stack and hardware-native DSL for frontier ML kernels, built around storage-first tensor layouts and reusable tile primitives.
https://t.co/V1yHdxlpCi
On NVIDIA B200, TIRx delivers up to ~1.08× over cuBLASLt on dense GEMM, outperforms DeepGEMM on all FP8 blockwise workloads with up to ~1.09× speedup, keeps FlashAttention-4 (FA4) typically within ~±2% of CuTeDSL, and remains competitive with cuBLASLt/FlashInfer on NVFP4 GEMM.
Through our past experiences building frontier ML kernels, megakernels, and agentic kernel systems, we kept seeing the same boundary problem: new operators and new hardware require new optimization strategies that often break old programming models or compiler passes.
TIRx builds on top of Apache TVM and moves toward a simple goal: let users and agents express the best-performing program, even for future hardware generations, while keeping the engineering effort for new kernels and new hardware as low as possible.
Excited to share XGrammar-2 has been accepted by #ACMCAIS 2026 @CAISconf! 🎉
⚡️ Up to 80x speedup
🛡️ Strict tool-calling correctness
🚀 Trusted by xAI in production systems
Join our presentation today at 4 PM at Bayshore Ballroom to see how we power critical agent workloads at scale.
🔗Blog: https://t.co/N0Tbl58Grf
🔗GitHub: https://t.co/lo4yScvfRN
🚀 The wait is over! Today at #MLSys, we'll give a talk to reveal the final results and present the awards for the FlashInfer AI GPU Competition! 🏆
I'll also introduce FlashInfer-Bench: an agent-oriented Benchmark Engine designed for production kernels.
Join us from 11:00 AM - 1:00 PM PT to see who takes the crown and learn more. Everyone is welcome to attend—see you there! ✨
🌐 Competition & Results: https://t.co/GS21eemEZv
💻 FlashInfer-Bench Benchmark Engine: https://t.co/rlzNUXJq5e
#FlashInfer #MLSys26 #AI #GPU
🚀Introducing Motus, the open-source agent infrastructure that learns in production.
Existing agent infra serves static agents: the harness, model, and workflow are fixed after deployment. But static agents degrade over time. The harness goes stale, new models go unincorporated, context drifts, and latency compounds.
Motus closes this gap by learning from every trace (failures, latency, cost, and task outcomes) and using those signals to continuously optimize agent harness, model orchestration, context memory, and end-to-end latency.
Early results: higher accuracy than any single frontier model at 2.3× lower cost (Terminal-Bench 2.0, SWE-bench Verified), with 52% lower latency and 45% better memory recall.
Open source under Apache 2.0. Works with any agent SDK. Deploy with one command.
https://t.co/C4u6JUzige
https://t.co/QIfKIikZQb
Super excited to share that I'll be joining @CarnegieMellon as a PhD student, working with @tqchenml and @ericxing!
It has been a wonderful journey at @uwcse@UWSyFi learning and building systems that power frontier AI in production. I want to express my sincerest gratitude to @ye_combinator@tqchenml@luisceze for all the opportunities and guidance along the way, and to many others at UW and CMU who have been hugely encouraging, supporting, and intellectually inspiring me. I wouldn't have made it this far without all of you.
Looking ahead, I'm eager to explore how AI-system co-design can advance the capabilities of both sides. On the system side, I believe better abstractions and verification signal design can enable the AI-driven cycle for system improvements. On the model side, I'm interested in how to enable models to perform well in long-horizon, sparse-goal tasks that require periodic knowledge consolidation, like doing system research itself.
Always happy to chat and collab!
Keep building 🤟
After 5 amazing years, I’m leaving the PyTorch team at Meta. I did my best work there and got to work with some of the smartest, most OSS pilled engineers in the industry. More soon on what’s next: still systems, still OSS (but not everything), a smaller team with a lot of GPUs
🚀 MLSys 2026 Contest - @nvidia Track is LIVE!
Registration is now open for the FlashInfer-Bench Challenge! Submit high-performance GPU kernels for cutting-edge LLM architectures on NVIDIA Blackwell GPUs.
Three Tracks
* MoE (Mixture of Experts)
* DSA (Deepseek Sparse Attention)
* GDN (Gated Delta Net)
Human experts AND AI agents welcome — evaluated separately. Let's see who builds the best kernels! 🤖
🎁 Prizes: Winners take home NVIDIA GPUs and are invited for presentation at MLSys 2026.
⚡ First 50 teams to register get free GPU credits from @modal - huge thanks for the sponsorship @charles_irl !
Whether you're a kernel wizard or building autonomous coding agents, we want to see what you've got.
🔗 Contest details: https://t.co/0ILK1D4Z9o
See you at MLSys 2026! 🔥
Together with the FlashInfer community, we built FlashInfer-Bench — a benchmark of real-world, AI system–driven GPU workloads — and, more importantly, an infrastructure and workflow to 0‑day ship AI‑generated kernels into production.
It's been a wonderful journey building FlashInfer-Bench, can't wait to continue the journey of self-optimizing AI systems with the kernel agent compiler!
🤔 Can AI optimize the systems it runs on?
🚀 Introducing FlashInfer-Bench, a workflow that makes AI systems self-improving with agents:
- Standardized signature for LLM serving kernels
- Implement kernels with your preferred language
- Benchmark them against real-world serving workloads
- Fastest kernels get day-0 integrated into production
First-class integration with FlashInfer, SGLang (@lmsysorg ), and vLLM (@vllm_project ) at launch🙌
Blog post: https://t.co/RG3oMO9brO
Leaderboard: https://t.co/R9T4jJJNd0
Live from the AI Infra Summit, co-located with #PyTorchCon — Tianqi Chen (@nvidia) explores how shared ML foundations can advance interoperability across compilers, libraries, DSLs, and frameworks, while unifying workloads across edge and cloud.
🔗 https://t.co/lLaazPPW2z
#AIInfraSummit #OpenSourceAI #AIInfrastructure
@WFS_MsWilsonBA There was a lot of discussion about time. The power-point made sense because it was told from a child perspective. The power point conveyed an informal form.
@WFS_MsWilsonBA The title A to B was explained by Bosco, which hints a theme, A signifies success and fame, B signifies losing and negativity. The book discusses a lot about shifting from positive to negative over time. Also, I’m in the zoom room.