Frontier pretraining is said to be a big-lab-only game. We don’t have 100k chips yet, so there’s only one way: algorithmic efficiency. Our new recipe matches DeepSeek V4 Pro’s pretrain using 50x less compute – that’s roughly half the FLOPs used for GPT3, or ~$0.5M on GB200.
https://t.co/uBs5o5Arsu
Super impressive. Congrats!
And maybe suggests that during RSI, the automated AI researchers will be less bottlenecked by compute than we might naively think.
Trtllmgen kernels are now open. Fastest prefill and decode kernels for our target workloads. We wrote these to win InferenceX, MLPerf, other benchmarks. Powering some of today’s top served models. Dive in, learn, use them, or level up your own. Enjoy.
https://t.co/2aQBwcdnZL
Sponsored ads: the AI colosseum is alive and well — Google’s still the gladiator ring where models battle for your clicks. The road to most tokens still runs through top clicks. ⚔️
#Sponsored#ads#OpenAl#Anthropic#Codex#Claude
@_sanjoydas and @clattner_llvm PTX generation isn’t just emiting one PTX instruction. Yes, the backend (.td it or inline it) must cover legacy + new PTX. But the game is (a) sequencing, vectorizing, and interleaving of PTX within a warp, and (b) choreographing across warps in a warp-specialized kernel on Hopper and more on Blackwell. I read these results as : Mojo matches and/or beat alternatives on (a) and (b).
Thank you to folks at @metaai for publishing their independent perf analysis comparing CUDA and Mojo against Triton and TileLang DSLs, showing Mojo meeting and beating CUDA, and leaving DSLs in the dust.
We just launched Mojo🔥 GPU Puzzles Edition 1, a hands-on guide that teaches GPU programming through 34 progressive challenges, not lectures.
Learn by doing, from your first GPU threads to tensor cores.
Works on NVIDIA, AMD, and Apple GPUs.
https://t.co/nXy5cDSJCU
Blackwell TMEM turns register spills from bad to a kernel design strategy. Need bigger tiles, larger faster memory. Spill/Fill RMEM to/from TMEM. Guess what the spilled data can be filled in other warps too :). #Blackwell#Kernels