Made with Claude Opus 5.5.
The fearmongering about AI always makes us forget that NOTHING WENT FOOM, as was always predicted.
Send this to your doomer friend who has a very high P(Doom).
Accelerate.
Made with Claude Opus 5.5.
The fearmongering about AI always makes us forget that NOTHING WENT FOOM, as was always predicted.
Send this to your doomer friend who has a very high P(Doom).
Accelerate.
@beffjezos Honestly? Fuck em’. You were thrown into a crazy situation and dealt with it best you knew how. Fighting for the future of humanity is no small task.
Training a 2B model from scratch on a 2080 Ti (11GB).
84% of it is a 1.7B engram table injected into the residual stream. 312M is transformer, 20M active per token. fused MoE, val 3.812 19h in, 18d left
round 1: teacher T0 generates a corpus of reasoning traces. Not just answers but full chains of thought, with the dead ends included. Student S1 trains on those traces. Standard distillation so far. Round 2: S1 becomes T1. It generates a new corpus. But here’s the beautiful catch, T1 generates traces that are cleaner than the ones it trained on, because during training it learned which parts of T0’s reasoning were strong and which were noise. The distillation acts as a filter. Signal goes up. Round 3 onward: this compounds. Each generation’s traces are cleaner than the last. The model is not learning new facts about the world. It is learning to think more efficiently about the facts it already has. Compression of cognition. Am I on the right track? 😉
New NanoGPT Speedrun WR at 68.0 (-5.8s) from @theonlyglitch_ , with a decrease of QK dim from 128 to 96, a kernel fusion of [QK Norm, RoPE, KeyOffset, paired head layout] into Triton, and moving MLP bwk, QKV fwd/bwk to FP8. This is a heavily involved PR with over 1k lines of triton, 4 different FP8 scaling protocols, and fancy register aware epilogue placement.
Two takeaways: 1) From a model perspective, if QK dims smaller than 128 works better at nano scale, then perhaps larger than 128 works better at Hero scale. 2) This Nvidia Engr is super legit. https://t.co/qBrhWUuvI4
I work at OpenAI and personally think AI has been and will continue to be an extremely beneficial technology to humanity
The conversation should be around how many billions of lives it will save
@mark_k Well would you look at that. I wouldn't be surprised in the least if it comes out that Sanders was involved. The timing? Coincidence? I think not.
@beffjezos This is truly awful. I have a feeling this guy is a huge burning Sanders supporter and I wouldn’t put it past him that he actually had a conversation with Bernie before pulling this stunt.