@adamheimann@wafer_ai@AMD scale up scale down, amortization across workloads, caching via l2/l3.
also use an moe instead of 27b dense for prefill speed. also dont use vllm for mamba if you want caching to work properly for those long prefills. sglang radix cache sir is better for this sir.
π¨ BREAKING:
these engineers figured out how to serve Kimi K3 on @AMD MI355X at 952 tok/s/node and 118 tok/s single stream!
this crushes B200 by 3.8x in aggregate throughput/node and 1.3x in single stream decode + beats B300 on performance per dollar (48 vs 33 tok/s/$)
See how in the thread.
my team figured out how to run Kimi K3 on @AMD MI355X at 952 tok/s/node and 118 tok/s single stream.
3.8x the aggregate throughput/node and 1.3x the single stream decode of B200 and beat B300 on performance per dollar: 48 vs 33 tok/s/$
details in reply