Do you like more speed? I know I do ⚡️
Run Qwen3.8-27B on your @NVIDIAAI DGX Spark with @sgl_project ✨
SGLang • up to 1M context • MTP
~33-35 tok/s single stream (structural)
~195-210 tok/s with 10 concurrent streams
No difference in speed when running with 256k or 1M context. Default is 256k to support 10 concurrencies with full kv cache.
Get it here:
https://t.co/rooS9ykL91