GPU from 2017 just pushed Qwen3.8-27B to 251 tok/s.
1Cat-vLLM 1.5.0 is out. 🐱
4× Tesla V100:
⚡ Qwen3.8 DFlash2: 206–251 tok/s
⚡ 256K decode: 50.38 tok/s
⚡ Long-prefill Attention: ~60.8 TFLOP/s
V100 isn’t getting newer.
The software is getting better.
Make Volta Fast Again.
https://t.co/WbrDoBGIUF