@elonmusk C based ML Training will become more mainstream for
larger models.. Believe Karpathy showed we can train a GPT-3 small for something like $20 using this approach..
@tbpn Does this have ability to mix data sources in desired ratios and I assume do Continued Pre Training from some standard checkpoint?
Also CPT will not be cheap.
IMO we are looking at hundreds of thousands if not million $ surprise bills here.
@AlexanderSpangh Inverse RL was a term I heard from one of my managers about 5 yrs ago - learn Rewards from Policy - we had decided it was better if we called it the FAFO approach..
@manas_muduli We need daily connections to Dubai / Abu Dhabi and Singapore and then BBI will be truly globally connected and other airports will become feasible.
@flavioAd If I remember, there was news about benchmark hacking for Llama4.
KimiK2 OTOH rolling out great Trillion plus parameter models (maybe using the muon optimizer that ensures stable training at scale)..
@basicprompts Yeah - paused a long time on this very point when reading up on SwiGlu Activation for my work.
It just works and we don't know the mathematical underpinnings why. 🤷