@rfleury@2Sexy4MyGPU It’s a Big Lebowski quote, the dude is a 60’s counter culture guy, to him The Eagles represent rock becoming mainstream and corporate
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
@daltonc@Bjorn_k@mwseibel Nice, really love these convos, I had no idea you were still doing them. Watched the one on slop vs craft, really great takes, gave me the same old vintage YC feeling.
We raised $250M to accelerate building AI's unified compute layer! 🔥 We’re now powering trillions of tokens, making AI workloads 4x faster 🚀 and 2.5x cheaper ⬇️ for our customers, and welcomed 10K’s of new developers 👩🏼💻. We're excited for the future!
We beat Nvidia’s cuBLAS kernels on B200s in 170 LOC.
Using zero CUDA. Just pure Mojo.
Here’s exactly how we went from 1% to 106% of Nvidia benchmark perf from scratch (with code) 👇🧵
Modular engineers are using Mojo with Nvidia Blackwell to make matrix multiplication faster than cuBLAS. In part 1 of our series, we explain what matmul is, why it’s fundamental to LLMs, and give a quick history from Ampere to Blackwell.
https://t.co/Hx5By6CS40
Part II of GPU Puzzles is here! Welcome to the detective work of GPU programming: debugging.🕵️ Puzzle 9 walks through the debugging workflow + 3 common issues, while Puzzle 10 teaches you to use @NVIDIA's compute-sanitizer to find and fix race conditions: https://t.co/PN3qxd7AqM
Inference costs too much. We’re fixing that! We've partnered with @sfcompute to launch the Large Scale Inference Batch API: up to 80% cost savings, 15+ leading models, and real-time GPU spot pricing. Learn how in our launch video: https://t.co/AukYI0Deld
Super excited to partner with @inworld_ai on their 20x cheaper, state-of-the-art text-to-speech model. Go try it now free! 🆓
Our technical achievements with @inworld_ai enabled the lowest latency, fastest TTS inference platform available on @NVIDIA B200 🚀 We'll be dropping a joint technical report soon! 🔜🔥
From a single (small!) binary, modular provides industry leading performance on AMD MI300/325 (up to 50% faster than VLLM 0.9!) runs with top speed on NVIDIA H100 and is previewing SotA Blackwell support. It’s also the best way to accel trad Python! 😘
Modular Platform 25.3 is here, and it's a major milestone for open source AI! 450k+ lines of Mojo & MAX code, including the full Mojo std lib, the MAX AI Kernels, & the MAX serving library. All open source in one of the largest kernel releases ever 🤯 https://t.co/3VIMZ3Dcok
@clattner_llvm@jack_clayto On May 2nd, @clattner and Modular's Senior Director of GenAI Abdul Dakkak take the @GPU_MODE YouTube livestream to talk all things Mojo, MAX, and GPU programming, including a new tile-based Mojo programming model: https://t.co/AjQ0KgDeOA
On April 24th, we're live from Modular headquarters with Next-Gen GPU Programming: Hands-on with Mojo + MAX, featuring @clattner_llvm and @jack_clayto: https://t.co/hpAtX1cPML
We heard there's 1,001 ways to write CUDA kernels in Python 😹 Let's set aside that nightmare while @clattner_llvm dissects Triton and its eDSL cousins in part 7 of Democratizing AI Compute.
https://t.co/cOfCrA77nS
Have you ever been curious about how to program GPUs and how they work? Check out how simple it can be with Mojo🔥: https://t.co/aSnceMXFgY
This is a true alternative to NVIDIA CUDA C++, giving you access to all the low-level GPU primitives you expect, with Python-like syntax and modern systems programming language features.
I'd love your feedback and suggestions for topics to cover in following chapters.
Nobody is going to hand roll GPU assembly because the ISA changes with every new architecture. There is no backwards compatibility. The magic is in the compiler and the CUDA compiler extensions are tightly coupled to each architecture. Modular/MLIR and OpenAI Triton are working on this.
MAX 24.6 is here and available for download today! Say hello to MAX GPU: the first vertically integrated gen AI serving stack, with SoTA performance on NVIDIA A100 and support for deployment across all major clouds 🤯 Don't miss the announcement: https://t.co/YL4ZnASslZ