Day 2️⃣ and it’s our turn on the @vllm_project track.
2 PM, Breakout 3: our MLOps Engineer, Ananth will present lessons from serving GLM-5.2 (@Zai_org ) & @Kimi_Moonshot K3 on GB300 & B300 clusters.
Catch the talk, then come say hi at our booth and grab some Finnish chocolate 🍫
Day one at @raydistributed and @vllm_project Summit ⚡ and we're on the schedule today
3 PM, Lightning Theater: Helen Zhao (@RedHat_AI) on scaling drafter training in real time across a full Verda GB300 rack with vLLM, Mooncake & Speculators.
Booth's open all day on the expo floor. Come around, grab some Finnish chocolate 🍫 and talk shop: GPUs, distributed training, scaling AI in production.
Meet us at @raydistributed and @vllm_project Summit on Aug 25-26 in San Francisco 🌁
Chat with us about cutting-edge AI infrastructure and optimizing distributed workloads. We’ll announce our talk soon!
Meet your future colleagues. We’re hiring in the US and beyond: https://t.co/kX1nNdJBNI
Love seeing this 😍
The @RedHat_AI team built their DSpark speculator (~4x faster Kimi-K3 decoding) on our GB300 and the results speak for themselves.
Speculative decoding is one of the optimization dimensions we actively work on in the AI Lab, so this is exactly the kind of work we love hosting 👇
~4x faster Kimi-K3 decoding.
Our new DSpark speculator takes single-stream interactivity from ~110 to ~435 tok/s/user on math reasoning, delivers ~3.5x the output throughput at matched interactivity under load, and thanks to sliding window attention (2048-token window across all 5 draft layers), our model holds steady acceptance out to 20K context across 10 LongBench domains.
Check it out! https://t.co/AGQlV61Y8U
More details in reply:
As a full-stack AI cloud and NVIDIA Cloud Partner, we operate every layer this depends on: power, cooling, networking, orchestration, engineering. Compute becomes infrastructure when it stays productive over time, and when someone does the work to keep it that way. That is the work we are committed to.
Verda welcomes NVIDIA's partnership with leading financial institutions to bring long-term institutional capital into AI infrastructure.
Access to capital now determines who can build at scale. Widening it matters for the AI labs and model owners we serve, and for Europe's capacity to build independently.
July, by the numbers: 3 features shipped, 4 events, 1 new CFO 👇
🔐 Audit Logs
📦 Storage in binary units (GiB/TiB)
👔 New CFO
📰 Forbes, Helsingin Sanomat & Sifted
🌍 RAISE Summit, DDN AI Data Summit, ICML & Ai4
Full digest → https://t.co/rsdNFqxBb4
Our CTO Arturs Polis joined “Orchestrating the agentic AI era,” discussing the infrastructure needed as AI systems coordinate increasingly complex models, tools, and services.
Great to be part of the conversation as we continue working with Arm on infrastructure built for the next wave of AI workloads.
Learn more about our partner with Arm: https://t.co/7We4XUrP07
AI infrastructure is being redesigned for a world where agents, not just models, are doing the work.
Verda recently joined @Arm in Cambridge to continue the conversation from two angles. 🧵
Melious is now serving GLM 5.2 and DeepSeek V4 Flash 0731 in production on Verda.
European AI that stands on its own: open-weight models through a single API, now running on our clusters.
More about the partnership: https://t.co/ST4fIALIJ9
"The coming years represent a critical opportunity for European companies to scale and establish themselves as global leaders in AI infrastructure.”
Read more: https://t.co/JYC1wSs1is
We're thrilled to welcome Mikko Einiö onboard as Chief Financial Officer. Mikko brings 20 years of investment banking, capital raising and corporate development experience to support us in our next phase of growth 🧵
Everyone assumed cheaper AI = less compute needed. Turns out it's the opposite. Jevons Paradox, live and in production.
@Forbes covered it this week, with our CTO Arturs Polis on the infrastructure side. The demand curve isn't slowing down.
Full article ⬇️ https://t.co/hOEaGZkSDr