43 days ago Marin kicked off the largest live-streamed pretraining run in history: a 535B MoE trained on 18T tokens. A scaling ladder was used to forecast eval loss along the entire trajectory, and the forecast was publicly pre-registered. Its halfway along, and despite being a 300x extrapolation, the run is within 0.3% of the forecast.
The run was originally planned to be 360B parameters. Until @ravwojdyla found a hardware implementation with custom kernels to grow to 535B parameters while also speeding up tokens per second. Learn about how he did it on the Open Athena blog: https://t.co/aOtxDCr1El.
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s!
This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it.
Specifically:
-(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time.
-Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used.
-Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2.
-Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step.
-Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params.
-Hand-rolled flash attention for 64 dim heads.
There are several additions that add accuracy too:
-(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15.
-(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application.
-A couple additional dynamic skip connections in the network.
The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large,
only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead.
https://t.co/Ycrzy6JFC3
As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://t.co/YEfM1VpOTi
The post is written in the perspective "research needs things (dynamic compute, network access, tools, downloads, internet, GUI, etc) and security tries to make those needs safe". A security first mindset is "security needs things, and then research works within that box to make things productive". The framing of what is a Need vs a Want is important.
hidden_dim 10 w/ vocab 4475 means you have a massive bottleneck (loss of information) on the backward pass when you go from logit space to hidden_dim space. Adding more neurons doesn't change the bottleneck. (there are values the MLPs could take on to fit the data, but the gradient may not find them)
There’s an enormous amount of open training data on Hugging Face. What does it take to make it work together 🤗?
For Marin’s 535B run, we built on 25T tokens from 152 datasets with licenses permitting training.
Here’s the work between downloading those and training a model 🧵
New NanoGPT Speedrun WR at 67.6s (-0.4s) from Jan Varho (jvarho on GitHub). Simple idea: mask impossible continuations. EG: "Hyperparameter" tokenizes to ["Hyper", "param", "eter"]. "eter" can never follow "Hyper" in the tokenized dataset since "Hypereter" tokenizes to ["H", "ype", "re", "ter"]. Yet, during multi-token prediction, "Hyper" learns to predict both "param" and "eter". Explicitly masking "eter" helps improve val loss, and during inference prevents a model from generating token pairs its never seen in training. https://t.co/VhTpNwzw4B
New NanoGPT Speedrun WR at 68.0 (-5.8s) from @theonlyglitch_ , with a decrease of QK dim from 128 to 96, a kernel fusion of [QK Norm, RoPE, KeyOffset, paired head layout] into Triton, and moving MLP bwk, QKV fwd/bwk to FP8. This is a heavily involved PR with over 1k lines of triton, 4 different FP8 scaling protocols, and fancy register aware epilogue placement.
Two takeaways: 1) From a model perspective, if QK dims smaller than 128 works better at nano scale, then perhaps larger than 128 works better at Hero scale. 2) This Nvidia Engr is super legit. https://t.co/qBrhWUuvI4
Was reading "Lost in Backpropagation: The LM Head is a Gradient Bottleneck" by @nthngdy and @yoavartzi the other day. Their ideas seemed very clever, but my ability to follow their intuition was lacking. So I decided to imagine "What would it be like to work as a neuron", to better process their ideas. The main audience when writing this was just myself. Sharing in case others find it a fun read.
A Day in the Life of a Neuron Associate at Grub Zoo.
Welcome to Grub Zoo. Proud home of 1000 animals. Each animal has its own designated 5 minute feeding window during the day where it eats one serving of grub. The feeding schedule looks something like:
-Feed the horse and the Giraffe from 4:00-4:05.
-Feed the rhino from 4:15-4:20.
-Feed the 8 alligators from 7:10-7:15.
Initially the Zoo ran off a simple lookup chart, where one could query a 5 minute window, and get returned a list of all animals that needed to be fed. Well, Bob from Marketing had a great (terrible) idea to replace our lookup table with AI. But instead of using AI, we could just become AI (wtf?). He explained that we’d each get some role in a mega human neural network, then use the power of ‘backpropagation’ to tell each other what to do. I think this is his promo project or something, not sure. His idea is that the optimal feeding window should depend on all sorts of properties of the weather, and once we get this learning machine running, it can plug in a bunch of these advanced features.
The first part of the plan was to designate 10 employees as feeders. Each feeder had to write down how they would divvy up any grub they were given between the 1,000 animals. One feeder, Jeremy, decided he’d give all grub he was given to the sloth. Weird guy. Another one, Anna, decided she’d split up the grub equally across all 1,000 animals. (How’s she gonna manage that, lol?).
My job was the supplier. Every 5 minutes I’d show up with some servings of grub. Then I’d have to divvy it up between the feeders, who’d then divvy it up between the animals based on their prearranged fractions. The weird part here is that I was the only one who was aware of the time of day- the feeders always stuck to their fractions.
Bob said that during each 5 minute interval we’d have two phases: an execution phase, and a feedback phase. Well for the first 5 minute execution phase I frankly had no idea what was going on, so I just gave an equal amount of grub to each of the 10 feeders, who then distributed it out. Somehow Anna was able to get grub out to all 1,000 animals, impressive.
Only the horse and the giraffe were supposed to get fed in this window, but the small amount of grub ended up spread out across all animals, and nobody was happy. I voted to drop Bob’s idea, but he figured we could salvage things. His insight was “Most windows only have a couple animals who need getting fed, so after everybody distributes their grub, we should apply some sort of bunching action towards the ones with the most grub. That way at least somebody gets a decent meal. Sometimes multiple animals need grub and we won’t be sure we got it right, so just do soft bunching.” He called it softmax.
We applied his technique, and sure enough one animal got a good meal, but it was the wrong animal! And the animals at Grub Zoo are extremely picky about when they get fed! Well, Bob said this was ok at first (what?), we just needed the feedback phase.
The feedback phase started with the animals. The ones that didn’t get fed but were supposed to, complained loudly. The feeders were the only ones who could observe and hear the animals. Each feeder would first go observe all 1000 animals to see which ones were complaining. Then they’d update their fractions slightly towards giving them more grub. Now if the supplier had given the feeder little or no grub that window, the feeder would take little effort in updating their fraction as it wouldn’t matter anyway. But if they had been given a large amount of grub, they’d take tremendous honor in updating their fraction quite heavily. But they had to be careful here! They were going to use that fraction for every time window, so they didn’t want to over-correct for one 5 minute interval.
The next phase of feedback was for me as the supplier. Each of the 10 feeders came and told me either ‘Man I coulda made a difference I had the fractions right just give me more grub’, or ‘I was off on this one, didn’t need grub here’. Then my job was to update ‘Ok for this 5 minute interval, give grub to feeder X’, because they told me their fraction is calibrated. I figured none of the feeders had it exactly right, but if I gave more grub to the feeders that seemed decent at the right times, their errors would average out and softmax would help bunch it up on the right animals.
To me the process felt pretty dumb, as the feeders were just like mindless cogs, trying to find the one right fraction. They could see all 1000 animals. But all I could see were 10 people shouting ‘give me more/less grub’. And somehow I was supposed to manage all the windows from that?
Over time I developed a massive lookup table. For each 5 minute interval, how much grub to give to each of the 10 feeders. Each feeder started to get more sophisticated specialization. For instance, one guy fed just the 8 alligators. I knew to only give him grub from 7:10-7:15, because that’s when he complained.
Well, Bob got wind that I was using a lookup table for each 5 minute window and freaked out. Said it wasn’t true AI. Said I was getting demoted to a neuron. 9 other people got hired as neurons too, working in parallel to me.
Now instead of a unique fraction for each window, I had a simpler role. I had to learn a fixed ratio across the 10 feeders, and a simple rule for when to switch on. Initially I just put the fraction on for the morning. And then tuned the fraction a bit to give grub to the feeders that were clamoring for more grub in the morning. I figured if I helped focus on some of the morning intervals, other ‘neuron associates’ could manage other windows of the day. I never got to see what these other neuron associates were up to, but I figured if I listened to the feeders I could help compensate for the gaps.
Well, we got some new funding, and Bob decided to hire some more ‘neuron associates’, so we are up to 20 now. Communication is becoming a real problem. We never talk to each other, and it’s getting kind of weird only listening to the 10 feeders and trying to compensate for what all my coworkers are doing.
Now Bob had the craziest idea yet. He decided to break up our 20 neuron associates into two 10 member squads: Alpha and Bravo. I got put on Bravo. During each interval I’d get to see both the time of day, and how much the Alpha Squad was planning to give to each feeder. I decided to take on a targeted role: IF the time was between 10-11am AND I saw that Alpha Squad was giving grub to feeder 2, then I’d give grub to feeder 1 and 2. I learned that in the morning whenever Alpha Squad gave grub to feeder 2, feeder 2 would still complain and want even more, and feeder 1 complained as well.
I noticed that sometimes Alpha Squad would slip up and not give grub to feeder 2, which meant I didn’t trigger on feeder 1 or 2. When this happened, I’d complain loudly. Hey Alpha! Give more grub to feeder 2! Problem is, team Alpha was getting overwhelmed during the feedback cycle. They had to listen to both the feeders directly, and every single member on squad Bravo. The problem got even worse when we doubled the size of Alpha and Bravo to 20 people each. That’s when Bob started hiring a new role, called ‘transporters’ and ‘messengers’.
There was one messenger and transporter for each feeder. Everyone only got to receive feedback from the messengers, and give grub to the transporters. During the execution phase, the transporters would collect the grub from squad Alpha, then I could make my decision based on if transporter 2 had grub. All of us in squad Bravo would then add on to the grub the transporters were carrying. During the feedback phase, I’d tell messenger 2 if I wanted more grub sent to feeder 2. This way, Alpha squad didn’t have to talk to both the feeders and Bravo squad. The messengers just added up the feedback from both the feeders and Bravo.
We ran this operation for several weeks, with Alpha and Bravo squads distributing grub to the transporters, the messengers giving feedback, and the feeders divvying it out and observing the animals. It seemed like the dumbest thing I’d ever seen! A lookup table could have done the same thing!
Well, that’s when Bob made a change that actually blew my mind. He called it the ‘magic of generalization’. The first thing he did was give all neuron associates access to the temperature, in addition to the time of day. Now in the old world, we’d have to manually make updates to all 1000 animals in the lookup table to figure out how they responded to temperature. In the new world, I only had to change my allocation to the 10 transporters. The feeders had already built fractions that grouped the animals into correlated sets.
I developed a new trigger condition. IF the time was between 10-11am AND transporter 2 had grub AND the temperature was above 60F, then give grub to transporter 1, 2, and 3. I came to learn that in very cold weather, these messengers didn’t complain as much.
Over time Bob added new features I could trigger on. Rain. Wind. Sunshine. It was starting to be quite a lot to handle. Bob always had to go around explaining to every neuron associate how this new measurement would work. Eventually Bob got pretty tired of this. He decided that only squad Alpha could see things like time, temperature, rain, wind, and sunshine. At first this seemed completely crazy to me. How could I know when to trigger, without these key features?
We concocted a pretty funny plan. I noticed that messenger 9 NEVER complained. My guess is that feeder 9 was just completely checked-out. Never gave any grub to anyone. Team Alpha had picked up on this. They also knew that whenever it was between 10-11am and above 60F, messenger 2 would complain to them. So they decided to give grub to transporter 9 whenever it got hot in the morning- not because transporter 9 needed the grub. But because team Bravo could then use that signal to make better decisions.
I developed a new heuristic: IF transporter 2 AND transporter 9 have grub, then we give grub to transporter 1,2,3. AND we also take that grub away from transporter 9 and add it to the pile, since feeder 9 never complains.
So that became my job. I’d wake up every morning, drive in to work as a neuron associate, then pay very close attention to which 10 transporters had grub, and how that correlated to the complaints I’d get from the 10 messengers. I’d update my fractions among the 10 transporters, and update my feedback to the 10 messengers.
Later that month my team name got updated to Squad Bravo-46. What the heck? Turns out Bob had really scaled the operation. We now had 50 Bravo squads in series. My job was always the same, talk to 10 messengers, and observe 10 transporters. I had no idea what measurements got fed into Alpha Squad. And honestly, I had no idea how many animals even lived at the Zoo anymore. I didn’t even know how many team members were on Bravo-46, since I never talked to them.
Alpha Squad was the only team that got to see the pure, raw measurements. And the 10 feeders were the only team that got to see the actual animals. In particular, it seemed like the feeders functioned like quite the bottleneck. I'd love to get to observe all 1,000 animals, instead of just 10 messengers and transporters.
Anyways, that’s my job. Hope the giraffe is doing alright.
It could certainly be done in theory, and our last run we repeated some data up to 7 times. Some factors influencing why 18.75T: For very sparse MoEs, compute optimal is much lower TPP. A rule of thumb we have found roughly holds is 60 tokens per active parameter, which puts us at about 13x overtrained. There are diminishing returns to more tokens. The reason we dont train at strictly 'compute optimal' is inference efficiency during RL and serving, as well as MFU considerations due to memory limits.
With that said, there are still gains from 13x overtrained through 100x overtrained and beyond. However due to finite compute, we chose this tradeoff. If you expected 10000x inference usage compared to training, such as google AI Overview, then you'd pick a very different point on that tradeoff (much higher TPP).
🚢 Marin 535B-A23B started training this week! As usual, the whole process is open.
Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow.
Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
At Marin we’ve found that scaling laws let us simulate the entire training trajectory. This became clear in our last 67B MoE run. The run hit the loss target within 1%, but perhaps more interestingly, arbitrary points during training could also be predicted to within 1%. This model is finishing up long context and will soon enter RL. Based on this finding, we’ve used a scaling ladder to simulate the training trajectory of the 535B MoE we kicked off yesterday. Will see how it goes. Now if someone can figure out how to fit a scaling law to each parameter we can just predict our final model and skip this training business.
Something else I've learned from this process is that getting a training run off the ground is way more than just fiddling with the architecture. Lot of herculean efforts from teammates to get trillions of tokens curated, get hardware running on a new GPU stack, improve MFU, and propose great ideas that were tested and integrated.
Given that this is our largest run to date, I anticipate lots of learning and adapting along the way. But thats part of the fun. Details: https://t.co/N7ZxcgvkKs
@ChrisJMcCormick@codingfisch Well I've at least done a couple runs on 8xH100 before for some speedrun thingy 😂. tho we are not advanced enough yet for your manual backprop
Given we had never done a GPU run, getting everything spun up on JAX GPU on H100s for testing and B200s for prod at scale. We are traditionally a JAX shop on TPU, but had GPUs available for this run. Some aspects felt closer to the show Cutthroat Kitchen, where you have to quickly improvise your recipe given new constraints on what empirically runs quickly and doesn't.
@surmenok Many people, myself included, have tried fp8 and failed on this repo. So thought it was cool someone was able to find a way to make it work. Scales to larger models too!
New NanoGPT Speedrun WR at 73.8s (-0.8s) from Mister-dev-oss, CerovazS, MarioPaerle, GabrieleCirillo, and crisostomi on GitHub. They added an incredibly sophisticated fp8 implementation on the MLP down proj fwd pass. 280 lines that include putting the activation quantization step into the prior kernel, delayed amax scaling, and fusing the weight quantization and transpose together.
https://t.co/2MA8QIQoTi
New NanoGPT Speedrun WR at 74.6s (−0.8s), by Jan Varho (jvarho on GitHub). The clever idea: if the target is the token " bananas", give partial credit for predicting " banana". More generally, during early training, add an auxiliary loss on the longest token-prefix of each target token. https://t.co/GNPpeiSgzG