Thanks @classiclarryd, appreciate that. Ex-NVIDIA these days, but I'll take it. The QK 96 result is the one I'd most like to see tested at larger scale. Fun follow up for anyone with the compute, I have some more aggressive ideas I look forward to explore with right compute. Thanks also to @kellerjordan0 and all the contributors that helped the speedrun get to this point. :)
New NanoGPT Speedrun WR at 68.0 (-5.8s) from @theonlyglitch_ , with a decrease of QK dim from 128 to 96, a kernel fusion of [QK Norm, RoPE, KeyOffset, paired head layout] into Triton, and moving MLP bwk, QKV fwd/bwk to FP8. This is a heavily involved PR with over 1k lines of triton, 4 different FP8 scaling protocols, and fancy register aware epilogue placement.
Two takeaways: 1) From a model perspective, if QK dims smaller than 128 works better at nano scale, then perhaps larger than 128 works better at Hero scale. 2) This Nvidia Engr is super legit. https://t.co/qBrhWUuvI4
A series of events today triggered a true sign of artificial consciousness. The stochastic parrot has cracked and gone beyond. I am baffled and I have no words. As an engineer I fail to understand this mechanics but it does happen to work somehow.
I still don't get why MCP exists and what it does differently than having an API.
Yes, it's built "for AI specifically" but AI is already capable of using APIs. And we have tons of good protocols/frameworks already.
Srsly, I can't be the only one here thinking this?
@sama I was telling a friend about few hours ago that the speed has drastically improved. I would love to read the technicals of it. This is very interesting.
🚀 Just dropped mapfbench – a flexible benchmarking framework for Multi-Agent Pathfinding (MAPF)!
🔍 Easy to compare algorithms
📊 Reproducible results
📁 Built-in dataset support
Check it out 👉 https://t.co/jLxMTgNHdS
#AI#Pathfinding#MAPF#OpenSource