I wrote a matmul kernel on B200 in pure CUDA/PTX that beats cuBLAS by 6% at M=N=K=8192.
Inspired by @gaunernst's blog on Blackwell instructions with benchmarking done on @modal.
Blog: https://t.co/nifzaMdPJS
Repo: https://t.co/EKjrnCrp0z
@MainzOnX Sounds great, will let you know when I release it. Also very excited about where this could go, especially in regard to finding deadlocks in multi-gpu programing.
I wrote a matmul kernel on B200 in pure CUDA/PTX that beats cuBLAS by 6% at M=N=K=8192.
Inspired by @gaunernst's blog on Blackwell instructions with benchmarking done on @modal.
Blog: https://t.co/nifzaMdPJS
Repo: https://t.co/EKjrnCrp0z