@LiamFedus Curious about the midtraining result; does the "midtraining improves RL scaling" figure in the blog include midtraining compute on the x-axis? Wondering if it outperforms RL given compute(mid + RL) = compute(RL only)?
These are trades, and it's natural to think about them in expected value terms, but very few people would take a trade with a 10% risk of blowing up their bankroll no matter how high its EV is. People that claim they would either have a very weird utility shape or are being disingenuous about the probabilities.
@AgustinLebron3 You can easily construct arguments where a 10% risk of extinction is acceptable if the other 90% of outcomes is prosperous enough, e.g. cure all diseases for all future lives. Utility is asymmetric, and I'm not saying I have this take, but many people do.
The total number of mosquito neurons in the world is equal to roughly 24 million GPT-4s. Relatedly, given 110 trillion mosquitoes, how many would you expect to have metacognated?
Recent (natural language -> music) products make me wonder about the feasibility of this space altogether. I'm doubtful that you can map local language space neighborhoods to local music space neighborhoods, which feels necessary for any intentional tool.
https://t.co/uka12UYYt5