Stanford's CS 153 syllabus is basically a rolodex:
guest lectures from Jensen Huang, Sam Altman, Satya Nadella, Andrej Karpathy, Ben Horowitz, and Garry Tan.
Same course, same semester. This is the actual gap between "a CS degree" and "a CS degree at Stanford."
A 92-year-old man with no computer and no analysts sat down in front of one MBA class and explained how he beat the stock market for forty-five years straight.
He did it for free.
Almost no one has watched it since.
His name was Walter Schloss. He started on Wall Street in 1934, at eighteen, in the middle of the Great Depression. A year later he took Benjamin Graham's Security Analysis course at Columbia. Graham hired him onto his own team in 1946. In 1955, Schloss started his own fund out of a one-room office, with nothing but Moody's manuals for research.
From 1956 to 2000, he compounded money at 15.3% a year against 11.5% for the market, over forty-five years, losing money in exactly two of them.
The lecture was arranged by a professor at the Ben Graham Centre for Value Investing at the Richard Ivey School of Business. No slides, no notes. Just a ninety-two-year-old man telling a room of MBA students how not to lose money.
His entire method fits in a handful of rules.
Buy stocks trading below tangible book value. Avoid companies carrying real debt. Never talk to management, because it only clouds your judgment. Hold fifteen to twenty positions at once instead of betting big on one. Selling is harder than buying, and the higher a stock climbs, the more it should scare you.
He repeated some version of one line roughly ten times in that lecture.
"Do not lose money."
Hedge funds today run teams of analysts and Bloomberg terminals chasing what Schloss did alone with paper manuals and one room.
Warren Buffett once singled him out by name as one of the greatest investors he'd ever known. Almost nobody else has ever heard of him.
Schloss died four years after this lecture, at ninety-five. Forty-five years in the market. Two losing years.
A $500+ course today. Stanford gave it away free back in 2018. 901,000 views โ and almost nobody stays until the part where Andrew Ng derives the sigmoid from scratch.
This isn't an accident. The ML course industry makes money on packaging: slides, animations, "intuitive explanations." The original is just a guy at a whiteboard, deriving the math live. That's exactly why most paid courses avoid showing the full derivation.
REASON 1: LOCALLY WEIGHTED REGRESSION
Instead of one global linear model, every new query point gets its own local model, weighted by distance to the training examples:
w = exp(-(x - xโ)ยฒ / 2ฯยฒ)
One parameter, ฯ, decides whether the model overfits or not.
It's non-parametric โ no fixed number of parameters, and it has to keep the entire training set in memory. Too small a ฯ and the model memorizes noise. Too large and it smooths out real signal. The whole overfitting/underfitting tradeoff fits into one letter.
Why does squared error even make sense? Assume errors are Gaussian noise ~ N(0, ฯยฒ). Multiply the likelihood of every data point together. Maximize that log-likelihood โ and you land on the exact same formula as "minimize sum of squared errors."
Least squares isn't an arbitrary convention. It's what falls out of probability theory once you write down the Gaussian assumption. Most courses just hand you the formula. Ng shows where it comes from.
REASON 2: LOGISTIC REGRESSION โ SAME TRICK, DIFFERENT PROBLEM
The sigmoid isn't glued onto linear regression for convenience. It falls out of the same likelihood question โ just for a binary outcome:
h(x) = P(y=1 | x; ฮธ) = 1 / (1 + e^(-ฮธแตx))
The log-likelihood here is concave โ a global maximum is guaranteed, nowhere to get stuck.
REASON 3: NEWTON'S METHOD
Instead of crawling toward the answer with gradient descent, Newton's method uses the Hessian matrix and almost jumps to the solution:
ฮธ := ฮธ - Hโปยนโฮธโ(ฮธ)
Convergence is quadratic โ the error roughly squares each step. The cost: inverting the Hessian gets expensive as the number of features grows.
The most uncomfortable takeaway from the whole lecture: what gets marketed as "advanced math"
MIT's Erik Demaine on beating O(Nยฒ):
the trick isn't a clever comparison, it's grouping. Split the sorted list into small blocks โ size log N or โ(log N) โ then compare pairs of blocks with a specialized lookup structure instead of searching finger by finger.
A $500+ course today. Stanford gave it away free back in 2018. 901,000 views โ and almost nobody stays until the part where Andrew Ng derives the sigmoid from scratch.
This isn't an accident. The ML course industry makes money on packaging: slides, animations, "intuitive explanations." The original is just a guy at a whiteboard, deriving the math live. That's exactly why most paid courses avoid showing the full derivation.
REASON 1: LOCALLY WEIGHTED REGRESSION
Instead of one global linear model, every new query point gets its own local model, weighted by distance to the training examples:
w = exp(-(x - xโ)ยฒ / 2ฯยฒ)
One parameter, ฯ, decides whether the model overfits or not.
It's non-parametric โ no fixed number of parameters, and it has to keep the entire training set in memory. Too small a ฯ and the model memorizes noise. Too large and it smooths out real signal. The whole overfitting/underfitting tradeoff fits into one letter.
Why does squared error even make sense? Assume errors are Gaussian noise ~ N(0, ฯยฒ). Multiply the likelihood of every data point together. Maximize that log-likelihood โ and you land on the exact same formula as "minimize sum of squared errors."
Least squares isn't an arbitrary convention. It's what falls out of probability theory once you write down the Gaussian assumption. Most courses just hand you the formula. Ng shows where it comes from.
REASON 2: LOGISTIC REGRESSION โ SAME TRICK, DIFFERENT PROBLEM
The sigmoid isn't glued onto linear regression for convenience. It falls out of the same likelihood question โ just for a binary outcome:
h(x) = P(y=1 | x; ฮธ) = 1 / (1 + e^(-ฮธแตx))
The log-likelihood here is concave โ a global maximum is guaranteed, nowhere to get stuck.
REASON 3: NEWTON'S METHOD
Instead of crawling toward the answer with gradient descent, Newton's method uses the Hessian matrix and almost jumps to the solution:
ฮธ := ฮธ - Hโปยนโฮธโ(ฮธ)
Convergence is quadratic โ the error roughly squares each step. The cost: inverting the Hessian gets expensive as the number of features grows.
The most uncomfortable takeaway from the whole lecture: what gets marketed as "advanced math"
MIT's Andrew Lo tells skeptical students the same thing every semester:
he's not trying to make you a finance major. He's trying to make sure you speak the one language every business function runs on โ because finance isn't a specialty, it's the lifeblood.
Insurance is the oldest of the four ways. It is a nine-trillion-dollar global industry. The equation underneath it was invented in 1560 by a broke Italian gambler.
His name was Girolamo Cardano. He wrote a book called Liber de Ludo Aleae. A short manual on how to win at dice. Nobody in finance read it for four hundred years.
Then in 1996 a ninety-year-old man in New York wrote a book that traced every modern risk model back to that manual. He called it Against the Gods. One thesis. Every dollar of premium ever collected on Earth is a footnote to a gambler scribbling in Milan.
His name was Peter Bernstein. He founded the Journal of Portfolio Management in 1974 and ran money at Bernstein-Macaulay before that. Wall Street called him the historian of risk.
In 2008 a small production company filmed him for thirteen minutes. He walked through the entire five-hundred-year arc. Cardano to Pascal to Fermat to Black-Scholes. Then he stopped and said the industry had built glass towers on the back of an idea a broke Italian scribbled to settle a card debt.
He died the following summer. Age ninety.
Reinsurance premiums crossed six hundred billion dollars last year. Every actuary on Earth prices catastrophe risk with the same expected-value framework Cardano invented to shave the house edge in Milan.
The video is thirteen minutes and twenty-two seconds long. Free. Eleven years on YouTube. Twenty-nine thousand people have watched it.
Almost none of them work in insurance.
A Yale physics professor's intro to quantum mechanics:
"The bad news is it's impossible to understand intuitively. The good news is nobody can." Even Feynman agreed.
His real goal: get everyone equally confused, then send them out to spread it further.
A finance professor tells every incoming class the same thing: you don't need to become a finance major.
You just need to understand it โ because finance is the lingua franca every business function eventually has to speak, whether you planned on it or not.
In August 2025, Stanford recorded a regular Friday lecture. No launch, no hype tweet.
It now has hundreds of thousands of views โ more than most viral AI explainers get in a year.
The lecturers are Afshine and Shervine Amidi. If you've ever bookmarked an ML cheat sheet, it was probably theirs.
The lecture doesn't start with attention mechanisms. It starts with why RNNs had to die first โ vanishing gradients, no memory of distant context, zero parallelization.
You can't actually understand self-attention until you've seen exactly what it was built to fix.
That's the whole lecture: not a definition of the Transformer, but the failure it was a direct answer to.
102 minutes. Free. On Stanford's YouTube channel.
Save this โ you'll want the tokenization section again later.
A model can score 100% on its training data and be completely useless. Most people building with AI right now don't know why.
Andrew Ng explains it in the very first lecture of Stanford's CS229 โ recorded in 2018, before transformers, before ChatGPT. It's still the clearest map I've seen for why any of this works.
Most explanations of ML start with algorithms. This one starts somewhere else โ and that's exactly why it holds up.
1.THE REAL SHIFT ISN'T THE ALGORITHM. IT'S WHERE THE RULES COME FROM.
Classical programming: Rules + Data โ Output. A person writes the logic.
Machine learning: Data + Examples โ Learning Algorithm โ Model. The system estimates the logic itself.
Try writing explicit rules for recognizing a face across every lighting condition, angle, and expression. You can't. That's not a limitation of programmers โ it's a limitation of rules. Machine learning exists because some problems are too messy for rules and clean enough for patterns.
2.THERE ARE ONLY THREE WAYS A SYSTEM CAN LEARN
Everything in ML collapses into one question: do you know the correct answer during training?
๐ Supervised Learning Know the answer? Yes. Goal: predict outputs. Example: spam detection.
๐ Unsupervised Learning Know the answer? No. Goal: discover structure. Example: customer clustering.
๐ฎ Reinforcement Learning Know the answer? No, but you get feedback. Goal: learn good decisions. Example: a game-playing agent.
Every ML system you've ever used fits into one of these three boxes. LLMs, recommendation engines, fraud detection โ all downstream of this split.
3. HERE'S WHY 100% ON TRAINING DATA MEANS NOTHING
Train a model, check accuracy: 100%. Looks perfect.
Run it on new data it's never seen: 55%.
It didn't learn the pattern. It memorized the noise. This is overfitting, and it's the single most common failure mode in machine learning โ a system that looks brilliant on the data you already have and useless on the data that actually matters.
The fix isn't a smarter algorithm. Ng's point, repeated across the whole course: the real skill isn't picking a fancier model. It's diagnosing why the current one fails โ bad labels, missing data, wrong features, high bias, high variance โ and running the experiment that actually answers the question.
Put together: rules give way to patterns, patterns split into three learning styles, and none of it matters if the model can't generalize past the examples it was shown. That's the entire discipline in one sentence.
Three questions worth sitting with:
1. Next time you evaluate any system โ a model, a hire, a strategy โ are you checking training performance or real-world performance?
2. Where in your work are you still writing manual rules for something that's actually a pattern-recognition problem?
3. When something underperforms, do you reach for a fancier solution before diagnosing what's actually broken?
Reply with your answer to just one of these. I'll read every one.
Save this before your next ML interview, course, or "wait, how does AI actually work" conversation. This is the one-paragraph version everyone wishes they'd gotten first.
Follow for the next one: how linear regression turns "find the right parameters" into actual math โ and why gradient descent is the method almost every model since 2018 still quietly depends on.
Dr. Jason Arday โ Cambridge's youngest Black professor โ was found dead days after resigning amid plagiarism allegations he denied. UK media ran 188 articles on him in 9 days. His family says he faced a "campaign of misinformation." Cambridge says it's investigating new claims.