🍰 another slice of CAKE just dropped: blackwell KDA kernels, ~2.05x geomean speedup vs FlashKDA:
https://t.co/hQBCTLUmXR by @avyyhuang
more cake kernels are baking 👀 not saying what CAKE actually is yet though, the full reveal is in the oven too.
Introducing XGrammar-2: structured generation for complex agent harnesses.
Strict tool-calling formats. Built-in DeepSeek-V4 and Qwen-3.6 support. Up to 80x speedup over XGrammar. Ready-to-use integrations with vLLM, SGLang, TensorRT-LLM, and more! ⚡
From Claude Code to OpenClaw, agents are defining more complex harnesses. XGrammar-2 ensures LLMs always interact with them in the right way.
Built in collaboration with DeepSeek, Databricks, and leading frontier AI labs to bring XGrammar-2 into latest models and products.
🧩 Structural Tag: one unified abstraction to describe any format your agent needs
🚀 Scales to 500+ strictly typed tools for complex agent harnesses
🌐 Native APIs in Python, C++, Rust, and JS, running everywhere from cloud to edge
🛠️ Integrated with vLLM, SGLang, TensorRT-LLM, and more
Excited to see what agent builders create with it!
Blog: https://t.co/N0Tbl588BH
GitHub: https://t.co/lo4yScuI2f
🚀Introducing Motus, the open-source agent infrastructure that learns in production.
Existing agent infra serves static agents: the harness, model, and workflow are fixed after deployment. But static agents degrade over time. The harness goes stale, new models go unincorporated, context drifts, and latency compounds.
Motus closes this gap by learning from every trace (failures, latency, cost, and task outcomes) and using those signals to continuously optimize agent harness, model orchestration, context memory, and end-to-end latency.
Early results: higher accuracy than any single frontier model at 2.3× lower cost (Terminal-Bench 2.0, SWE-bench Verified), with 52% lower latency and 45% better memory recall.
Open source under Apache 2.0. Works with any agent SDK. Deploy with one command.
https://t.co/C4u6JUzige
https://t.co/QIfKIikZQb
🚀 MLSys 2026 Contest - @nvidia Track is LIVE!
Registration is now open for the FlashInfer-Bench Challenge! Submit high-performance GPU kernels for cutting-edge LLM architectures on NVIDIA Blackwell GPUs.
Three Tracks
* MoE (Mixture of Experts)
* DSA (Deepseek Sparse Attention)
* GDN (Gated Delta Net)
Human experts AND AI agents welcome — evaluated separately. Let's see who builds the best kernels! 🤖
🎁 Prizes: Winners take home NVIDIA GPUs and are invited for presentation at MLSys 2026.
⚡ First 50 teams to register get free GPU credits from @modal - huge thanks for the sponsorship @charles_irl !
Whether you're a kernel wizard or building autonomous coding agents, we want to see what you've got.
🔗 Contest details: https://t.co/0ILK1D4Z9o
See you at MLSys 2026! 🔥
#MLSys2026 is inviting self-nominations for the External Review Committee (ERC)!
If you want to contribute to the review process for the MLSys conference, nominate yourself and help shape this year's program. We especially welcome PhD students and early-career researchers!
https://t.co/IPzuKOdbad
📢Excited to introduce Apache TVM FFI, an open ABI and FFI for ML systems, enabling compilers, libraries, DSLs, and frameworks to naturally interop with each other. Ship one library across pytorch, jax, cupy etc and runnable across python, c++, rust https://t.co/m2gHJRreol
🤔 Can AI optimize the systems it runs on?
🚀 Introducing FlashInfer-Bench, a workflow that makes AI systems self-improving with agents:
- Standardized signature for LLM serving kernels
- Implement kernels with your preferred language
- Benchmark them against real-world serving workloads
- Fastest kernels get day-0 integrated into production
First-class integration with FlashInfer, SGLang (@lmsysorg ), and vLLM (@vllm_project ) at launch🙌
Blog post: https://t.co/RG3oMO9brO
Leaderboard: https://t.co/R9T4jJJNd0