Is memory the moat?
@wafer_ai engineers figured out how to serve Kimi K3 on AMD MI355X at 952 tok/s/node and 118 tok/s single stream.
This is over 3.8x the aggregate throughput/node and over 1.3x the single stream decode of B200.a
This also crushes B300 on performance per dollar: 48 vs 33 tok/s/$.
@ianye0212 wrote about how we did this https://t.co/9uthIKPpMg
🚨 BREAKING:
wafer now runs the fastest Kimi K3 in the world
ranked #1 across every provider on @ArtificialAnlys
⚡ 172 output tok/s
⚡ 15.8s end-to-end response time
get this performance on a dedicated endpoint: https://t.co/TsEoUs3wpn
we're shipping kimi k3 FAST next week. fastest serverless kimi k3.
unlock 2x credits for kimi k3 Fast.
follow @wafer_ai and sign up here: https://t.co/rfurdSXFqF
🚨 BREAKING:
GLM-5.2 Fast (⚡️150-250+ tok/s) is 30%+ off through July 31.
old: $3.00 in / $10.25 out / $0.50 cache
NEW: $2.10 in / $6.60 out / $0.21 cache
live on @wafer_ai, @OpenRouter, and @vercel AI Gateway. pick wafer as your provider.
(screenshot from Vercel AI Gateway)