For those that hope (or worry) that LLMs will do breakthrough scientific research, I've got good (or bad) news:
LLMs are particularly, exceedingly, marvellously ill-suited to this task. (if you're a researcher, you'll have noticed this already)
Here's why🧵
New 3h31m video on YouTube:
"Deep Dive into LLMs like ChatGPT"
This is a general audience deep dive into the Large Language Model (LLM) AI technology that powers ChatGPT and related products. It is covers the full training stack of how the models are developed, along with mental models of how to think about their "psychology", and how to get the best use them in practical applications.
We cover all the major stages:
1. pretraining: data, tokenization, Transformer neural network I/O and internals, inference, GPT-2 training example, Llama 3.1 base inference examples
2. supervised finetuning: conversations data, "LLM Psychology": hallucinations, tool use, knowledge/working memory, knowledge of self, models need tokens to think, spelling, jagged intelligence
3. reinforcement learning: practice makes perfect, DeepSeek-R1, AlphaGo, RLHF.
I designed this video for the "general audience" track of my videos, which I believe are accessible to most people, even without technical background. It should give you an intuitive understanding of the full training pipeline of LLMs like ChatGPT, with many examples along the way, and maybe some ways of thinking around current capabilities, where we are, and what's coming.
(Also, I have one "Intro to LLMs" video already from ~year ago, but that is just a re-recording of a random talk, so I wanted to loop around and do a lot more comprehensive version of this topic. They can still be combined, as the talk goes a lot deeper into other topics, e.g. LLM OS and LLM Security)
Hope it's fun & useful!
https://t.co/75mXcUBI8L
@rpoghaziabad@passportsevamea The passport renewal application for Umang Gupta, with file number GZ1078709139623 was submitted on 19/10/2023, and unfortunately, it has been under review since then. Could you provide an update on when can we receive the passport?
Workin on slides to teach #machinelearning this Fall for @usfca_msds. Here is (1D) visual difference between L1 and L2 regularizations' effect on training loss function. Pushes loss up and min parameter (beta) towards the origin. That's why betas are constrained.
I'm working on a book on machine learning interviews so I've been spending the last few months talking to companies about their hiring process for ML roles. This thread is a summary of what I've learned. It will be updated as the book progresses. (1/n)
My father (52) died in an accident which happened due pure negligence of NHAI, contractors and Road Ministry. Refer the news attached @nitin_gadkari
@NHAISocialmedia @PMOIndia
Punjab: Three dead in Patiala highway mishap https://t.co/vaiqCsWKdE via @timesofindia
https://t.co/ypT44YkLpn
Here, I tried to explain building blocks of SOTA ULMFIT model. What is an AWD-LSTM? How Dropout is used everywhere? What is a QRNN and why might it be better? ...I also used excel spreadsheets to simplify things in a different way :)
Interested in the visualization of machine learning algorithms? Come Check out my talk in @DataInstituteSF / @usfca_msds seminar series, this coming Friday Nov 30 at 12:30 at @usfca downtown . Gonna be a hoot! https://t.co/8XQfNuMLRn
Using GANs to generate Master[Finger]Prints that unlock 22-78% phones sensors (dep. on security level of sensor) https://t.co/KeBLOv2H58 .. doesn't get much more "adversarial" than that.
Wanna know more about our @usfca_msds program or @DataInstituteSF at University of San Francisco? See interview with Director David Uminsky https://t.co/iFMeWXsyEN
Hi everyone,
Learn how to find underlying topics within a large text corpus: An introductory tutorial on topic modeling
https://t.co/33hTMXOeI5
#topicmodeling#Unsupervised#MachineLearning#LDA
To understand the methods used to build recommendations systems and the metrics for evaluating their effectiveness, head over to my latest blog post!
https://t.co/yL4J5rRJZP
#DataScience#recommendations#MachineLearning#DeepLearning