A few weeks ago we stood up our own multi-rack K3 deployment and I told everyone @inferact to dogfood it.
Nobody went back since then. It's better and faster at most of what we do daily, not to mention it never blocks us mid-task.
This is the power of open intelligence.
On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop.
Thank you to @AMD and @inferact for co-hosting! More meetups to come soon.
πππShoutout to @SemiAnalysis_ for this tremendous effort on benchmarking real-world agentic workloads!
Building an efficient, performant and reliable inference system for real frontier agentic workloads is what we care about and stay laser-focused on at @vllm_project and @inferact. We don't ship features that look good on paper or benchmax on 8k/1k - we optimize for what matters on real workloads - TTFT, interactivity, and cache hit rate!
Check out the released results where vLLM delivers SOTA performance across normalized interactivity and cost efficiency!
vLLM Conference is next week, and we have a packed schedule π π . Here's the full list of events you should know:
Mon:
π· 4β6PM vLLM Γ Ray Γ Google Cloud Happy Hour: https://t.co/Exw10SVIrV
π· 6β9PM vLLM Γ Dynamo Meetup: https://t.co/id6TaOfIMs
Tue:
π· 11β11:30AM vLLM Keynote from @simon_mo_
π· 12β5PM vLLM Track Day 1
π· 6β9PM vLLM Γ AMD Happy Hour: https://t.co/RVXTxGq0Lz
Wed:
π· 12β5PM vLLM Track Day 2
π· 6β8:30PM vLLM x DigitalOcean Γ NVIDIA Happy Hour: https://t.co/U1FSK6aYhZ
No ticket needed for the happy hours and meetups, but space is limited. To join the full event, register here: https://t.co/bTDMkOOwIR
The @vllm_project maintainers at @inferact π are some of the most cracked engineers in the world. Theyβre building one of the inference engines that powers much of the worldβs intelligenceβand doing so with remarkable dedication, kindness, and hard work.
With @inferact, @vllm_project is now officially verified by @Kimi_Moonshot to serve Kimi K3 with full accuracy, according to the Kimi Vendor Verifier.
As the leading inference engine, vLLM delivers best results across vision, 1M context, and agentic coding; with prod quality.
The vLLM Conference is coming up in 3 weeks! π
Come learn about the current state and future of AI inference, Aug 24β26 in San Francisco π, hosted by @inferact at @anyscalecompute Ray Summit.
We'll have speakers from Inferact, NVIDIA, AMD, Google TPU, Anyscale, PyTorch, Meta, Red Hat, and key builders around vLLM. The talks on the roadmap deep dive into the latest on accelerators, training and serving pipelines, and production-scale inference π
The most widely used open-source model inference engine @vllm_project built their own testing dashboard + an overnight bot to find what broke and draft reverts on top of our platform data
vLLM hit new peak bs=1 decode on Kimi-K3: 464 tok/s π
Under a low-entropy reasoning workload, Kimi-K3 + DSpark on vLLM reaches 464 tok/s on batch size 1 with 4Γ4 GB300.
This benchmark is fully reproducible with public image: vllm/vllm-openai:kimi-k3 and @inferact's DSpark draft model linked in the thread.
1/3
With Kimi K3 Day-0 on vLLM: Open Frontier Intelligence for Everyone π
At 2.8 trillion parameters, Moonshot AI's Kimi K3 is one of the most powerful open-weight models ever released. Starting today, you can serve it on vLLM the moment the weights are public.
What K3 brings:
π§ 2.8T-parameter Mixture-of-Experts (16 of 896 experts active per token)
π 1M-token context window
π οΈ Native multimodal understanding, including vision
β‘ Kimi Delta Attention: a hybrid of linear and full attention that makes million-token context affordable
Huge thank you to @Kimi_Moonshot AI for the model release and partnership, @inferact for leading the vLLM optimizations, and to our partners at @nvidia, @AMD, and the broader vLLM community.
1/6
@vllm_project The best part of this is that whenever vLLM posts numbers, they are real, not inflated, and not cherry-picked. That's actually quite rare in the industry. If vLLM gives us a number, we can trust it's reproducible and real.
With Kimi K3 Day-0 on vLLM: Open Frontier Intelligence for Everyone π
At 2.8 trillion parameters, Moonshot AI's Kimi K3 is one of the most powerful open-weight models ever released. Starting today, you can serve it on vLLM the moment the weights are public.
What K3 brings:
π§ 2.8T-parameter Mixture-of-Experts (16 of 896 experts active per token)
π 1M-token context window
π οΈ Native multimodal understanding, including vision
β‘ Kimi Delta Attention: a hybrid of linear and full attention that makes million-token context affordable
Huge thank you to @Kimi_Moonshot AI for the model release and partnership, @inferact for leading the vLLM optimizations, and to our partners at @nvidia, @AMD, and the broader vLLM community.
1/6
Great writeup from @khluu000 π
How does vLLM stay production-quality while merging ~2,000 commits/month and shipping every 2 weeks?
The team broke down the three layers that make it possible. A huge community effortβthank you to everyone who is helping along the way.
https://t.co/F9FoltezVp