vLLM hit new peak bs=1 decode on Kimi-K3: 464 tok/s ๐
Under a low-entropy reasoning workload, Kimi-K3 + DSpark on vLLM reaches 464 tok/s on batch size 1 with 4ร4 GB300.
This benchmark is fully reproducible with public image: vllm/vllm-openai:kimi-k3 and @inferact's DSpark draft model linked in the thread.
1/3
Kimi K3 license. It's inspired by MIT but distinctly non-commercial, where any company making over $20M/yr must get a specific commercial deal (and display Kimi K3 if over 100M users or $20M/mo revenue)
For my first post, Iโm sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
Feels bigger than the DeepSeek moment. Its clear that benchmarks don't tell the whole story, and that Fable/Sol tokens are likely still higher ROI on many tasks, but still - this is crazy.
Kimi K3 is the best performing model on https://t.co/aporqgIfIh, ahead of Fable, reaching a comparable success rate in less time.
This is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark.
Notes:
โช๏ธ Benchmarks donโt always tell the full story, although this is important signal, adding to mounting evidence that this could be a breakthrough moment for open models
โช๏ธ No model as of yet has reached 100% completion on this set of evals. The top performer peaks at 92% and 96% โwith helpโ
@Kimi_Moonshot Congrats guys on the release and thanks for supporting opensource! ๐๐ฅฐ
We'll hopefully be able to make 1-bit GGUF quants so people can run them locally (if you can) ๐ญ
@0xBADB01E OSS models provide an equalizer here, increasingly donโt have to just rent to frontier labs to turn compute to $ + compute ratio of inference:training will keep increasing.