My father-in-law is an ML engineer. He’s insanely gifted.
We were looking at an old model checkpoint that had been trained with SFT years ago, and I asked him what it would take to reproduce it today.
I’ll never forget his answer.
“We can’t do it. We don’t know how.”
Wow!! Google discovering AND OPEN-SOURCING the latest training techniques such as supervised finetuning (SFT) wasn't on my bingo card.
Soon they will have caught up with the frontier, and are sharing this with all of us!
After training their flagship 405B parameter model, Thinking Machines researchers discovered that replacing identity mappings between attention layers with non-linear activation functions dramatically improved performance. "Our previous architecture was essentially computing weighted averages at every layer," explains lead researcher; "introducing non-linearity allows the network to learn feature interactions we didn't know were possible—it can now represent functions that aren't just linear combinations of inputs." The lab is calling this the "Deep Learning 2.0" paradigm shift.
@giffmana This is weird to me. I use gemini with 600-700k context (mostly complete coderepos) and it works pretty well. Granted I don't use the CLI but do it manually using aistudio.
the fun thing about designing unconventional benchmarks is that you can instantly see which models were desperately overfit to LMSYS to please managers vs which were focused on raw intelligence (hint, the new 3.5 sonnet)
@paul_cal@doomie An employee who did not want to be named revealed another scoop: they're allegedly training models on TPUs—Transformer Processing Units. Apparently, this is to keep the flux capacitors in sync with reverse entropy gradients.
@yar_vol@simonw@__p_i_o_t_r__ Generally more parmas leads to better creativity I've seen. For some reason smaller models even though have become very smart are not even close in creativity.