@giffmana@Emy_Aze@JFPuget@RheaSukthanker@CameronPashmina That is insane work calling Lucas Beyer a bum. Bro damn near reinvented computer vision. If he read the title of my paper AND gave me feedback, Iโm putting that on my CV.
@thinkymachines Why spend more time post-training an older architecture for agentic performance when you have an all star team that can build a better arch and better training paradigm.
@LiamFedus@jietang So continuous pretraining is pseudo-continuous learning and post-training is what mobilizes parameters into useful units of reasoning?
@jietang I am witnessing this firsthand with one of my PhD projects, the pre-training holds immense value, especially as you trace and follow where the model is underperforming. I accidentally left my model pretraining using a different sampling method, now I have hope Iโll graduate.
@sama This will probably save quite a bit of money on training, hence diverting more resources to inference and perfecting Codex. Maybe even potentially building out other products/projects/side quests.