Top Tweets for #MLsys
vLLM v0.30.0 freezes gc during CUDA graph capture now. H200: capture time 12s to 2s, engine init 28.9s to 8.2s. one gc.freeze() call clawing back 20 seconds every cold start.
python garbage collector was quietly taxing your GPU boot time
#vLLM #Inference #MLSys
SGLang v0.5.21 lets you flip a prefill server into a decode server with one curl. No restart, no weight reload, KV pool stays put. 100 role switches in their stress test, 0 fails. But who decides when to flip ? Your router.
#LLM #AIInference #MLSys

vLLM before v0.31.0: a LoRA named "foo" and a cache_salt of "foo" hash to the same prefix-cache block for the same prompt. Ran the key logic, collide: True. v0.31.0 shipped today and tags every key by source.
#LLM #AIInference #MLSys

Big congratulations to U-M CSE PhD student @jaewon_chung_cs on being named a 2026 @MLCommons Rising Star! ๐
Recognized for his impactful research in making AI systems more energy-efficient. ๐
๐Read more: https://t.co/dtc5sNq3kT
#UMichCSE #MLSys #GreenAI #MachineLearning

We were happy to sponsor #MLSys 2026. Across the talks, posters, and keynotes, three themes defined the current state of inference serving:
1. Agentic engineering
2. KV Cache optimization
3. Heterogenous hardware
Our read on each: https://t.co/GWRa8YKaFH
๐ The wait is over! Today at #MLSys, we'll give a talk to reveal the final results and present the awards for the FlashInfer AI GPU Competition! ๐
I'll also introduce FlashInfer-Bench: an agent-oriented Benchmark Engine designed for production kernels.
Join us from 11:00 AM - 1:00 PM PT to see who takes the crown and learn more. Everyone is welcome to attendโsee you there! โจ
๐ Competition & Results: https://t.co/GS21eemEZv
๐ป FlashInfer-Bench Benchmark Engine: https://t.co/rlzNUXJq5e
#FlashInfer #MLSys26 #AI #GPU

8/ Congrats on the leadership, Zhengding Hu @hu_ding17008 Yufei Ding @creamyoki Chao Zhang @chaozhangcs from #UCSD and #GaTech
#LLM #AIAgents #RSI #MLSys
Iโm in MLSysโ26 Wednesday morning presenting our work CAD on disaggregated LLM training for long context
If youโre interested in talking about LLM training and inference infra in general, come and letโs chat!
#mlsys
๐ฅCAD: Efficient Long-context Language Model Training by Core Attention Disaggregation
Repo: https://t.co/QdNk8iXy6c
Blog: https://t.co/O5xRrl22UJ
Training a long-context LLM model can suffer from severe workload imbalance caused by core-attention - the softmax(QK^T)V part.
Core-attention disaggregation (CAD) fundamentally eliminates workload imbalance by disaggregating core-attention from the rest of the model.
Is anyone attending #MLSys next week? Would love to connect!
Iโll be attending and am based in Redmond, so Iโm happy to answer questions or share Seattle-area recs if youโre visiting for the conference.
#MLSys2026

Iโll be at #MLSys this week, May 18โ22 ๐
PyTorch Foundation will have a booth with experts on PyTorch, vLLM, Ray + other foundation projects. Come by, ask questions, and meet the teams building open AI infra ๐ฅ
Iโm also speaking Monday morning on agentic self-improvement with OpenRoll ๐ค
See you there ๐
#PyTorch #vLLM #Ray @PyTorch @vllm_project @raydistributed @linuxfoundation @aaif_io

Our paper #ExecuTorch - A Unified PyTorch Solution to Run #ML Models #OnDevice is accepted to appear at #MLSys 2026. Excited!
Exciting work from our team, studying data efficiency for RLVR. These kinds of insights inform our dataset creation work for foundation model labs. Kudos to @realjustinbauer @pham_derek for this paper's acceptance to #MLSys 2026!
Our paper โLearning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimesโ was accepted to #MLSys 2026!
We introduce three procedurally generated, verifiable datasetsโCounting, Graph, and Spatial Reasoningโto study RLVR under low-data / low-compute constraints.
Key result: small, mixed-complexity datasets can be more data-efficient than large, easy ones.
Our paper โLearning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimesโ was accepted to #MLSys 2026!
We introduce three procedurally generated, verifiable datasetsโCounting, Graph, and Spatial Reasoningโto study RLVR under low-data / low-compute constraints.
Key result: small, mixed-complexity datasets can be more data-efficient than large, easy ones.
๐ Thrilled to share our #MLSYS 2026 #FlashInferBench competition with @nvidia!
Write blazing-fast GPU Kernels with your AI agents on the latest Blackwell GPUs, and go head-to-head with strong human kernels in #FlashInfer.
๐ Register by Feb 15
๐ https://t.co/V6g31eoLVf
๐ MLSys 2026 Contest - @nvidia Track is LIVE!
Registration is now open for the FlashInfer-Bench Challenge! Submit high-performance GPU kernels for cutting-edge LLM architectures on NVIDIA Blackwell GPUs.
Three Tracks
* MoE (Mixture of Experts)
* DSA (Deepseek Sparse Attention)
* GDN (Gated Delta Net)
Human experts AND AI agents welcome โ evaluated separately. Let's see who builds the best kernels! ๐ค
๐ Prizes: Winners take home NVIDIA GPUs and are invited for presentation at MLSys 2026.
โก First 50 teams to register get free GPU credits from @modal - huge thanks for the sponsorship @charles_irl !
Whether you're a kernel wizard or building autonomous coding agents, we want to see what you've got.
๐ Contest details: https://t.co/0ILK1D4Z9o
See you at MLSys 2026! ๐ฅ
We have arrived in St. Louis for @Supercomputing 2025. Learn more about our group's research through the various talks, panels and tutorials below. We will also be at the @UofMaryland booth (3123) at various times. #SC25 #UMDCS #HPC #AI #MLSys #AI4Science

๐Excited to share the #MLSys Call for Papers!
For the first time, weโre also welcoming submissions to the Industrial Track.
Research and industrial track deadline: Oct 30, 2025
Reviews available: Jan 12, 2026
Author responses: Jan 16, 2026
Notifications: Jan 25, 2026
https://t.co/yArzEAwQDc
https://t.co/o3hikWB14o
Calling industry researchers: MLSys 2026 launches its first Industrial Track! ๐
We're excited to announce the inaugural Call for Industrial Track Papers at MLSys 2026! ๐
๐ https://t.co/0vr2DjlwGb)
This is a unique opportunity for industry researchers and practitioners to share real-world innovations, system deployments, large-scale ML challenges, and lessons learned from practice with the MLSys community.
๐ Details & submission info:
Paper submission deadline: Oct 30, 2025 20:00 UTC
Full CFP: https://t.co/NpcmfDq19k
Iโm honored to help launch this new track and look forward to seeing your contributions that bridge cutting-edge research with impactful practice.
#MLSys2026 #CFP #MLSystems #MLforSystems
Last Seen Hashtags on Sotwe
Most Popular Users

Elon Musk 
@elonmusk
241.5M followers

Barack Obama 
@barackobama
118.9M followers

Cristiano Ronaldo 
@cristiano
114.6M followers

Donald J. Trump 
@realdonaldtrump
111.8M followers

Narendra Modi 
@narendramodi
107.1M followers

Rihanna 
@rihanna
98.7M followers

NASA 
@nasa
92.3M followers

Justin Bieber 
@justinbieber
91.8M followers

KATY PERRY 
@katyperry
90M followers

Taylor Swift 
@taylorswift13
84M followers

Lady Gaga 
@ladygaga
75.5M followers

Virat Kohli 
@imvkohli
73.5M followers

Kim Kardashian 
@kimkardashian
70.9M followers

YouTube 
@youtube
68.8M followers

Neymar Jr 
@neymarjr
66.5M followers

Bill Gates 
@billgates
65.2M followers

Selena Gomez 
@selenagomez
63.1M followers

The Ellen Show
@theellenshow
62.2M followers

CNN 
@cnn
61.8M followers

X 
@x
60.7M followers



















