@Freyabuilds Zuckerberg has a point: Apple’s biggest challenge may not be the iPhone itself, but what comes after it. Staying on top requires constant innovation, and history shows that no tech giant stays untouchable forever.
A 92-year-old man with no computer and no analysts sat down in front of one MBA class and explained how he beat the stock market for forty-five years straight.
He did it for free.
Almost no one has watched it since.
His name was Walter Schloss. He started on Wall Street in 1934, at eighteen, in the middle of the Great Depression. A year later he took Benjamin Graham's Security Analysis course at Columbia. Graham hired him onto his own team in 1946. In 1955, Schloss started his own fund out of a one-room office, with nothing but Moody's manuals for research.
From 1956 to 2000, he compounded money at 15.3% a year against 11.5% for the market, over forty-five years, losing money in exactly two of them.
The lecture was arranged by a professor at the Ben Graham Centre for Value Investing at the Richard Ivey School of Business. No slides, no notes. Just a ninety-two-year-old man telling a room of MBA students how not to lose money.
His entire method fits in a handful of rules.
Buy stocks trading below tangible book value. Avoid companies carrying real debt. Never talk to management, because it only clouds your judgment. Hold fifteen to twenty positions at once instead of betting big on one. Selling is harder than buying, and the higher a stock climbs, the more it should scare you.
He repeated some version of one line roughly ten times in that lecture.
"Do not lose money."
Hedge funds today run teams of analysts and Bloomberg terminals chasing what Schloss did alone with paper manuals and one room.
Warren Buffett once singled him out by name as one of the greatest investors he'd ever known. Almost nobody else has ever heard of him.
Schloss died four years after this lecture, at ninety-five. Forty-five years in the market. Two losing years.
@MadCrash_X Honestly, this is such a powerful reminder that great investing doesn’t have to be complicated. Schloss focused on the basics, stayed disciplined, and let time do the work. “Do not lose money” says it all
@MadCrash_X This is the kind of course that reminds you the math is usually less scary than the way it’s packaged. Seeing the derivations from scratch makes the whole thing click differently.
Tom Holland plays Spider-Man in front of a billion-dollar audience.
Off camera, he says he genuinely dislikes Hollywood — and calls the industry "scary."
A $500+ course today. Stanford gave it away free back in 2018. 901,000 views — and almost nobody stays until the part where Andrew Ng derives the sigmoid from scratch.
This isn't an accident. The ML course industry makes money on packaging: slides, animations, "intuitive explanations." The original is just a guy at a whiteboard, deriving the math live. That's exactly why most paid courses avoid showing the full derivation.
REASON 1: LOCALLY WEIGHTED REGRESSION
Instead of one global linear model, every new query point gets its own local model, weighted by distance to the training examples:
w = exp(-(x - x₀)² / 2τ²)
One parameter, τ, decides whether the model overfits or not.
It's non-parametric — no fixed number of parameters, and it has to keep the entire training set in memory. Too small a τ and the model memorizes noise. Too large and it smooths out real signal. The whole overfitting/underfitting tradeoff fits into one letter.
Why does squared error even make sense? Assume errors are Gaussian noise ~ N(0, σ²). Multiply the likelihood of every data point together. Maximize that log-likelihood — and you land on the exact same formula as "minimize sum of squared errors."
Least squares isn't an arbitrary convention. It's what falls out of probability theory once you write down the Gaussian assumption. Most courses just hand you the formula. Ng shows where it comes from.
REASON 2: LOGISTIC REGRESSION — SAME TRICK, DIFFERENT PROBLEM
The sigmoid isn't glued onto linear regression for convenience. It falls out of the same likelihood question — just for a binary outcome:
h(x) = P(y=1 | x; θ) = 1 / (1 + e^(-θᵀx))
The log-likelihood here is concave — a global maximum is guaranteed, nowhere to get stuck.
REASON 3: NEWTON'S METHOD
Instead of crawling toward the answer with gradient descent, Newton's method uses the Hessian matrix and almost jumps to the solution:
θ := θ - H⁻¹∇θℓ(θ)
Convergence is quadratic — the error roughly squares each step. The cost: inverting the Hessian gets expensive as the number of features grows.
The most uncomfortable takeaway from the whole lecture: what gets marketed as "advanced math"
MrBeast says he could start a brand-new channel today — no face, no voice, zero promotion — and hit 20 million subscribers in six months.
He says it's not luck. It's a formula.