Speedrunning GPT-2 is now routine thanks to @karpathy.
But can we speedrun GPT3-175B?
We attempted to match accuracy on a <$10K budget; while we didn't quite reach it, our first results show that quality data, engineering, and native FP4 can get close.
Details in 🧵
What’s the best model you can train in a day if someone hands you a pile of Blackwell GPUs?
You can try out yourself
On April 9 in Paris, @GPU_MODE + @verdacloud + @sestercegroup are hosting a GPU hackathon with a bunch of GPUs to run on and even more of them for the winners.
🚀 We are releasing state-of-the-art post-training quantization (PTQ) algorithms for Microscaling FP4, together with kernels:
- First study focused on MXFP4/NVFP4 PTQ for LLMs
- New Micro-Rotated (MR) format and GPTQ algorithm
- QuTLASS GPU kernels with up to 3.6x speedups.
❗️ We just expanded our capacity of B200 SXM6 180GB servers – available in our Cloud Platform.
The best thing is…
With DataCrunch, you can deploy the Blackwell platform without approvals.
Just sign in and select the instance type:
https://t.co/psjrOd88cX
The paper also suggests Group Tied Attention (GTA), which works in the opposite direction and draws inspiration from MLA, incorporating those techniques into GQA.
First of all, a confession! In the blog titled 'Multi-Head Latent Attention: Benefits in Memory and Computation', we didn't tell the whole story—the benchmarking on a single GPU. In reality, for DeepSeek V3-style models, parallelization is needed.
🆕 Inference API for FLUX.1 Kontext [max] & [pro] are now available on DataCrunch!
We are an infrastructure partner of @bfl_ml for Kontext, a suite of generative flow matching models for text-to-image and image-to-image editing.
Learn more: https://t.co/BTgTKAhI0g