OpenAI's Noam Brown says that while AI model performance scales roughly equivalently with more training or inference compute, the cost of inference is on the order of 100 billion times cheaper
Check out our latest Netflix techblog post about how we innovate on the reward function for our recommendation models to try and improve long-term member satisfaction. https://t.co/6TSmHeTqbQ
I gave a talk at Seoul National University.
I titled the talk “Large Language Models (in 2023)”. This was an ambitious attempt to summarize our exploding field.
Video: https://t.co/vumzAtUvBl
Slides: https://t.co/IidLe4JfrC
Trying to summarize the field forced me to think about what really matters in the field. While scaling undeniably stands out, its far-reaching implications are more nuanced. I share my thoughts on scaling from three angles:
1) Change in perspective is necessary because some abilities only emerge at a certain scale. Even if some abilities don’t work with the current generation LLMs, we should not claim that it doesn’t work. Rather, we should think it doesn’t work yet. Once larger models are available many conclusions change.
This also means that some conclusions from the past are invalidated and we need to constantly unlearn intuitions built on top of such ideas.
2) From first-principles, scaling up the Transformer amounts to efficiently doing matrix multiplications with many, many machines. I see many researchers in the field of LLM who are not familiar with how scaling is actually done. This section is targeted for technical audiences who want to understand what it means to train large models.
3) I talk about what we should think about for further scaling (think 10000x GPT-4 scale). To me scaling isn’t just doing the same thing with more machines. It entails finding the inductive bias that is the bottleneck in further scaling.
I believe that the maximum likelihood objective function is the bottleneck in achieving the scale of 10000x GPT-4 level. Learning the objective function with an expressive neural net is the next paradigm that is a lot more scalable. With the compute cost going down exponentially, scalable methods eventually win. Don’t compete with that.
In all of these sections, I strive to describe everything from first-principles. In an extremely fast moving field like LLM, no one can keep up. I believe that understanding the core ideas by deriving from first-principles is the only scalable approach.
In this blog post we share some of best practices and lessons that we learned while operating large-scale recommendation systems at Netflix.
https://t.co/cuMrGF0bER
#recommendersystems#mlops#personalization
It was my please to talk about RecSysOps in Nvidia RecSys Summit. The talk covers some of the lessons and best practices for running a large scale recommender system. The video of the talk is now available https://t.co/9ZlyPWlUBh
with @JustinBasilico#recsysops#recsys
Want to know how RecSysOps helped @NetflixResearch reduce production issues, increase #recsys quality, and decrease time spent on debugging issues? Join @ehsan_saberian from Netflix, on July 28 at our Recommender Systems Summit. https://t.co/PcE6v5ZTuL
Just making sure everyone read “The Bitter Lesson”, as it is one of the best compact pieces of insight into nature of progress in AI. Good habit to keep checking ideas on whether they pass the bitter lesson gut check https://t.co/xf6UwaqohF
Over the last couple of years, I've spent 1000's of hours building ML models.
Truth is, after using dozens of models/architectures, 99% of them are a waste of time.
I start most problems with 1 of 6 architectures.
Here are the best models for a strong baseline 🧵
I spent 500+ hours on Kaggle competitions last year and just became a Kaggle Master.
Over those many hours, I learned a systematic process you can use to train any model on any dataset.
6 steps to train any model 🧵
Transformers are arguably the most impactful deep learning architecture from the last 5 yrs.
In the next few threads, we’ll cover multi-head attention, GPT and BERT, Vision Transformer, and write these out in code. This thread → understanding multi-head attention.
1/n
We're happy to release #CausalML v0.12. It surpassed 637K downloads & 2.5K stars on GitHub. Big thanks to the team & all the contributors including 4 new community contributors. Please check out the release note: https://t.co/u6CH8AwYNM @UberOpenSource@paullo0106@totteh
We have extended the deadline for the workshop in Personalization and Recommendations in Search (PaRiS ) at WSDM 2022 @WSDMSocial
New submission deadline is Jan 14th.
We are excited to announce the #ICML 2022 Call for Tutorials!
We welcome proposals for tutorials on core machine learning topics and on topics of emerging importance for machine learning. Please see the call for further details👇
https://t.co/izMzjyd09H
@icmlconf
Check out the workshop on Personalization and Recommendations in Search coorganized by our @__sudarshan__ and @moumita_bh at #WSDM2022. Submissions due by 12/18.
How to share your progress with your mentors/collaborators?
Throughout your research project, 99% of the time your approach DOESN'T WORK (yet). 😬
How could we share these "failed results" and have productive conversations with your mentors/collaborators? 👇
Excited to co-organize a workshop on Personalization and Recommendations in Search with @__sudarshan__, Anlie Dong, @feifeiM, to be held at #WSDM2022. We invite submissions of original as well as preliminary research: https://t.co/ubruNbcWWK
#IR#personalized#MachineLearning
Interested in building the systems that we use to create the machine learning algorithms for personalizing the Netflix homepage? We're hiring a software engineer to join our team. Read more and apply here: https://t.co/ebMkBZV9JU