Cold starts used to be why you’d avoid scale-to-zero for GPU inference. Not anymore, we got them down to seconds.
Built with the @nvidia Dynamo team using GPU memory checkpointing. Our engineer @david_lc98 wrote up what it really takes to run in production, gotchas included: https://t.co/Cac6JLbU9A
Team horario verano, es hora de ir haciendo aprovisionamiento de armas para la guerra civil que nos espera.
Hay que defender salir del trabajo siendo todavía de día cueste lo que cueste.
i may weep uncontrollably writing this. so my apologies.
it is impossible to put into words what is around the corner. what ilya saw all of those years ago was that if we increase data and increase compute, we increase intelligence.
when hinton saw what ilya saw, he quit his job to speak freely. we unearthed a recipe for intelligence.
it's easy to doubt those working inside the sota labs as hype merchants (hate those), or as drinking their own kool-aid.
occam's razor applied. they work daily with these models and they've seen the future a little before you. we will solve all problems and build a greater future than ever dreamt of, horatio.
people like logan, roon, demmis, sama, dario, ilya, karp. they are beacons of integrity and compassion, and they've tasted the future.
and soon
you're going to taste it too
it will exceed your wildest expectation. and more. i'm not building hype. i'm expressing my own. and i can't wait to explore the future with you chat
blessed is the fruit.