@inductionheads@InverseMarcus hm? the SP paper was written in late 2020 and published in 2021 while the original chatgpt was released nov 2022. agreed that RL methods are ancient and that parrot claims are now obsolete, but the parrot claims in 2020/2021 were not made about chatgpt.
@inductionheads@InverseMarcus but the training did not always use that harness, e.g., original seq2seq paper (as well as many others obviously). it's not required. using a decoding harness during training is a distinction one can make when describing methods.
@InverseMarcus hard to answer w/o a definition of an LLM, as some think about them as a model family & others as systems with specific types of training criteria (string prediction). the latter may say current things are LLMs in terms of model family but not training.
@yoavgo@petergostev agreed. when we discuss the SP paper, we should note that it comments only on "systems which are trained on string prediction tasks", which would exclude most RL as used today and many other things.
@mathelirium@DavidPe51482177 but aren't you making some assumptions about the variance in those many dimensions? what if variance is near-zero in nearly all dimensions?
@jsuarez@andreiC_dev what are you maximizing when automatically tuning the coefficients? is it still some sort of automated metric/reward? if so, then why not just use that metric/reward directly?
@tinkerteller@yoavgo how about universal transformers (https://t.co/bsZyWaOAwI)? they can "can choose the number of processing steps based on the input data"
@ayaannmalik@Kangwook_Lee this output is a little misleading though.. the problem is that the o3 bar's height was erroneously set to equal the height of the 4o bar, not that the "without thinking" bar was drawn taller than the o3 bar.
@orionweller gotcha thanks! yea makes sense if you're considering baselines that are good for both types of tasks. i haven't used encs much for retrieval, but for classification i am still waiting for something to outperform debertav3-large. modernbert also did not outperform it.
@skirano correct answers implies supervised data created with manual effort (or synthetic data generated by templates.. still requires manual effort). i just don't see how this feels any different from the requirements of supervised finetuning.
@skirano huh? SFT still helps though. DeepSeek-R1-Zero is not as good as DeepSeek-R1. and DeepSeek-R1-Zero still requires labeled training data in order to compute rewards. the feedback you mention is correct answers.
@sharifshameem@abhi_venigalla "sometimes ~2 sec"? what does that mean exactly? for the 10x to be accurate, the mean or median would have to be 2sec, right?