No More GIL!
the Python team has officially accepted the proposal.
Congrats @colesbury on his multi-year brilliant effort to remove the GIL, and a heartfelt thanks to the Python Steering Council and Core team for a thoughtful plan to make this a reality.
https://t.co/58QK2yctRD
Doing a back-of-the-envelope calculation, a 7B Llama 2 model costs about $760,000 to pretrain!
And this assumes you get everything right from the get-go: no hparam tuning, no debugging, no crashes and restarting. The real cost, including experimentation, is probably millions!
It makes me appreciate all these openly available models!
Math:
- The total number of GPU hours needed is 184,320 hours.
- The cost of running one A100 instance per hour is approximately $33.
- Each instance has 8 A100 GPUs.
That's 184320 / 8 * 33 = $760,000
@Avishai_Bitton I agree. @aboutido and I do this every day. We’re just developing a GTM and execute with no prior expertise. We both learn constantly on task we’ve never done before. We left our confirt zone in fundraising daily.