@yunta_tsai Andrej Karpathy, OpenAI founding member, explains why the technique training every frontier model is "terrible" and the one thing all of them are still missing. "Reinforcement learning is a lot worse than I think the average person thinks. Reinforcement learning is terrible.
@LindaTangUSA While the cost per token (i.e., the compute needed for each token) has decreased, many of the latest models are significantly larger. This means that the overall compute required for a complete task or inference session can still be high, even if each token is processed cheaply.