If you find parity/phasebits in kernel warp specialization confusing I made some colorful illustrations here: https://t.co/hOIzsksLok
Wrote this while extremely sleep-deprived so apologies for any typos
Had a great time meeting all the smart folks at ICML🇰🇷, and thanks to everyone who stopped by the ParallelKernelBench poster. If you're at all interested in collaborating, please please reach out!
Special thanks to @togethercompute for supporting this project the whole way!
Multi-GPU kernels are the real test for coding models. Today at @aiDotEngineer, @simran_s_arora shared ParallelKernelBench, an open-source benchmark for evaluating whether LLMs can write fast CUDA kernels for real communication-heavy workloads.
Proud to see this work from the Together AI Frontier Performance team.
Struggling to write Ring Attention on TPUs/GPUs with @khshind was one of the original motivations for KernelBench 😅
It feels full circle with ParallelKernelBench — a dedicated eval to see whether LLMs can write fast multi-GPU kernels 📡
Introducing the latest KernelBench family member: PKB, led by awesome undergrad researchers @opengroundsFX & @NathanPaek9368! (+ the always amazing @simran_s_arora@realDanFu for their guidance 🙏)
Excited to release PKB Parallel Kernel Bench, led by Willy Chan and Nathan Paek @asplencmnt!!
A benchmark of mostly net-new multi-GPU kernel problems (solutions are independently useful for real-world workloads).
Excited to release ParallelKernelBench (PKB), a benchmark for measuring LLMs’ ability to write fast multi-GPU kernels! 😀
Multi-GPU kernel generation compounds several hard problems:
- a large parallelism design space
- a new communication axis to optimize
- and hardware-specific decisions around communication mechanisms
Existing kernel-generation benchmarks mostly target single-GPU workloads, so we built PKB to cover real-world multi-GPU workloads (many of which do not have existing optimized solutions). 🧵👇
LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart.
ParallelKernelBench measures how they fail by benchmarking against 87 problems pulled from real codebases including Megatron-LM, DeepSpeed, DeepEP, TensorRT-LLM, NeMo-RL.
New research from Willy Chan @asplencmnt@simonguozirui@simran_s_arora and @realDanFu
Modal is great to work with! Highly recommend if you're a researcher experimenting with GPUs
It definitely made adding new DSLs like Thunderkittens and TLX to the existing kernelbench infra a lot easier because container environments are really intuitive to specify and work with
Fresh blog post!
@modal partnered with @ScalingIntelLab, @HazyResearch, and @chelseabfinn's IRIS Lab to speed up research on speeding up AI research.
Read how scientists at the cutting edge are building the machines that build the machines with Modal.
https://t.co/FDbnp8DOuC