In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
I met @ssokota 5 years ago at a conference poster session. I was impressed with his explanation, so I offered him an internship. We continued collaborating and he now works with me at @OpenAI. Later, he told me I was the only person to stop by his poster for the entire session.
I really enjoyed working on this project. I believe the new flop-aware paradigm of nanoGPT thinking is useful, and it seems helpful in a compute optimization context like this one. It seems like a simple and logical conclusion of optimization: figure out what’s actually needed, and use that. The process was actually fun.
One note regarding this record, as noted in the blog. Some of the techniques that scale to frontier pretraining have been redacted from this record. Most importantly, ANVIL III, the significantly improved version of the ANVIL optimizer lineage, has not been included in this record, as it is proprietary to us at @hyperstition_cc. It replicates Muon’s speed but beats it by 20-28 millinats of loss at all model sizes 124M-1.2B and training depths (up to 8x-Chinchilla), with remarkably consistent hyperparameters and extremely high SNR, as compared to Muon. You can learn more about ANVIL III and Feather-1.7B, the model we trained using internal architectural changes and ANVIL III, here:
https://t.co/PyfvnSSwVC
Using ANVIL III (which allowed me to cut steps) improved the speedrun results by another 0.7 seconds, with lower net loss. This alone proves there’s still more room to improve this record.
One final thing that I’ve had many people ask about is how much of a role autoresearch played in my work. The answer is almost none. While almost all of the implementation was done with Claude Code, I did the ideation almost fully by myself. I spoke about this extensively in a conversation with @METR_Evals.
We had a cluster of H100s while I was working on the record. At night or when I stopped working on it, I’d send out several agents to try to hill-climb the speedrun and improve the record. In total, the agents made shockingly little progress, despite spending more time in aggregate than I did. No matter what I told them, it seemed difficult to get them out of local minimums, stuck performing futile tuning of a method that saved 0.1 seconds and cost ~5 millinats (clearly a bad trade), or just tuning hyperparameters for hours on end trying to make some small change. I saw similar phenomena as those described in this METR report:
https://t.co/ZtKPp2w7nT.
However, as I also discussed with METR, I don’t think the agents lack the knowledge to do this. Instead, I think a lot of it comes from a personality issue with current LLMs. They have very little belief that large jumps like this are possible, or they approach it with the wrong paradigm. Autoresearch was not very helpful in my work, and, from personal experience, I don’t think this current generation of publicly available models, without special harnesses, will do very high-quality autoresearch. However, I think the capability and knowledge is already here, which implies that we could be significantly closer to true autoresearch than recent results would imply.
NanoGPT was a really fun challenge and it led to a number of valuable insights and processes. Finally, I’d like to give @Classiclarryd and @KellerJordan a huge thank you for all of their efforts maintaining the repo.
strategy stealing is non-constructive in hex, even in odd-sized boards symmetric over both diagonals. (this margin is too narrow to contain my mediocre proof of this fact)
@jack_merullo_ For a token in this chart (e.g. "truths"), the landscape associated with it is the loss for predicting that token ("truths"), or for predicting the next one ("to")?
@N8Programs But the single step gradient is always calculated at the last step of each "stage", if I understand correctly. Seems weird that the gradient updates don't take the non-last steps into account at all.