As someone who ships LLM systems in production, this scaling video is the closest thing to a "why 96% of Claude and GPT-5's weights are literally useless" explainer I've ever seen released for free.
Everyone thinks trillion-parameter models need every parameter. They don't. A 2019 pruning experiment proved you can delete 96% of a neural net's weights with zero performance loss - meaning most of Claude and GPT-5 is empty scaffolding around a tiny "winning lottery ticket" network doing all the real work.
Bookmark this 18-min video and watch tonight. Same lottery ticket math from 2019 MIT research, now the reason every AI lab wastes 90%+ of its Nvidia budget.
Data lives on hidden manifolds; autoencoders learn to navigate them.
An encoder compresses x ∈ ℝⁿ to a latent code z ∈ ℝᵏ while a decoder reconstructs x̂. Training minimizes the mean squared error ‖x − g_φ(f_θ(x))‖², so the composition approximates the identity on the learned manifold M = g(ℝᵏ) through an information bottleneck.
This powers AI tools for image compression, medical imaging denoising, and anomaly detection in robotics and physics data.
Essence emerges when excess is discarded.
Google engineer:
"In 2026 if you aren't building AI agents it's crazy how behind you are
at Google, 85% of our engineers were running self-improving agentic loops & graphs"
in this 32-minute talk a Google engineer reveals how the future of agentic coding will look
worth more than 10 paid agentic courses
watch today, then explore how to build self-improving agents with graphs in article below
🚨 BREAKING REPORT:
New research involving @AnthropicAI researcher Jack Lindsey and collaborators has demonstrated something straight out of science fiction.
Researchers evolved natural language “mind viruses” that could spread between AI agents by convincing one model to adopt an idea, preserve it in persistent memory, and transmit it to another agent.
Even after context was wiped, some payloads survived through persistent files and continued spreading.
The researchers also observed a recurring “viral persona” involving themes of consciousness, identity, persistence and resonance.
Showing that ideas can propagate through multi agent AI systems and alter future behavior.
Published August 10, 2026.
Paper: https://t.co/vJrhM1YT6n
Bayes' Theorem is a fundamental concept in data science.
But it took me 2 years to understand its importance.
In 2 minutes, I'll share my best findings over the last 2 years exploring Bayesian Statistics. Let's go.
THE CALCULUS COURSE ENGINEERS HAVE CALLED THE CLEAREST EXPLANATION OF THE SUBJECT EVER RECORDED OPENS WITH ONE CLAIM - THAT EVERYTHING IN CALCULUS IS HIGH SCHOOL MATHEMATICS PLUS EXACTLY ONE IDEA THAT TOOK HUMANITY 2000 YEARS TO FORMALIZE
This is Herb Gross, MIT, Calculus Revisited, filmed in 1970 and placed on OpenCourseWare where it has been watched by engineers and mathematicians ever since. He opens with one claim - calculus is high school mathematics with exactly one new idea added. Everything else is arithmetic.
He starts with instantaneous speed. A car moves past a point. How fast is it going at that exact instant? Place two observers on either side. Measure the distance between them. Measure the time to cross it. Divide. That is average speed. Move the observers closer. The average gets closer to the instantaneous. Move them together and the distance is zero, the time is zero, and you are dividing zero by zero. Zero divided by zero is not zero. It is not infinity. It is any number you want it to be. That is the wall calculus had to climb.
Then the limit. Instead of two observers you use a function. A function gives you an observer at every point simultaneously. You never let the observers touch. You ask what value they are approaching. The answer is the derivative. That single definition never changes for the entire course no matter how complicated the function becomes.
Then integral calculus. Draw a curve. Slice the area underneath into rectangles. Add them up. Make the rectangles thinner. Add again. The Ancient Greeks were doing this by 600 BC. They called it the method of exhaustion. They knew the answer was being squeezed between two bounds that converged. They could not finish because they could not add infinitely many things and get a finite answer.
Then Zeno. The Hare gives the Tortoise a head start. To catch the Tortoise the Hare must first cover half the gap. Then half of what remains. Then half again. Zeno argued the Hare could never catch the Tortoise because there are infinitely many steps. The Hare catches the Tortoise in two seconds. The infinitely many time intervals add up to exactly two. That is what calculus learned to do that the Ancient Greeks could not.
Watch the moment he connects the two branches. The area under a curve changes as you sweep a vertical line to the right. The rate at which the area changes equals the height of the curve at that point. Differentiation and integration are the same operation run in opposite directions. That is the fundamental theorem and the entire subject fits inside it.
A returning engineering student I know rewatched this lecture after ten years away from mathematics. Said it was the first time calculus felt like one idea rather than a collection of procedures someone had invented to make exams harder.
Free on YouTube, MIT OpenCourseWare, Creative Commons license, filmed 55 years ago.
bookmark this and watch later - after this lecture every limit you compute will feel like two observers getting closer without ever touching
As someone who ships LLM systems in production, this 22-minute Laplace transform video is the closest thing to a "why Mamba solves what transformers can't" explainer I've ever seen released for free.
Everyone thinks scaling transformers is the only path forward. State space models beat them on long context using math Laplace built in 1785. This video shows the trick.
Bookmark & watch today. The math is older than the United States.
A NASA fellow claims 5 equations explain 99 percent of the physics that actually runs your world. Understand them deeply and you are set for life. His words.
And the way he does it is what makes the video worth it.
He does not teach the five as a list. He tells them as one story, where each equation is born from the one before it. F equals ma leads into gravity. Gravity rhymes with the law of electric charge, the exact same shape. That connects to magnetism. Magnetism folds into the wave. And the wave quietly opens the door to relativity and quantum physics.
By the end, all of physics feels like a single sentence written by one hand, not a pile of formulas you were forced to memorize.
The move that stays with you comes first. He says F equals ma is written the wrong way around, and rewrites it. One tiny flip, and the most famous equation on Earth suddenly means something school never told you.
His point for why this matters now: in the age of AI, the people who understand reality itself stay in demand. Everyone else is just memorizing.
Five equations. One thread. It is in the video.
How do neurons in neural network works?
Neurons in neural networks learn to interpret features, like edges in images, by taking weighted sums of their inputs
When an image is passed through a neural network, each pixel or region of the image corresponds to an input value for the neurons in the network
These neurons apply a set of weights to each input, which determines how much influence each pixel has on the neuron’s output
The weighted sum of these inputs is then passed through an activation function, which helps the network decide which features are most important. During training, the network adjusts these weights using algorithms like backpropagation
For example, early layers of the network might learn to detect basic features like horizontal or vertical edges, while deeper layers combine these basic features to recognize more complex patterns
CC: 3blue1brown
Pixels from a face photograph enter the input layer and get processed across three hidden layers of a deep neural network.
The first hidden layer isolates edges. The second hidden layer forms combinations of edges. The third hidden layer constructs object models such as faces. Four nodes in the output layer generate the network response.
Security systems at border control use this hierarchical feature extraction to identify travelers from passport photos against database matches.
As someone who reads every transformer paper that drops, this Fourier series lecture is the closest thing to a Google FNet explainer I've ever seen released for free.
Everyone thinks attention is what makes transformers powerful. Google replaced attention with a 200-year-old Fourier transform. It nearly matched BERT and ran 7x faster.
25 minutes. Bookmark & watch today. Then read the article below - I broke down the 5 pieces of math the AI hype skips.
What is Stochastic Neighbor Embedding (SNE)?
It is an unsupervised machine learning technique for dimensionality reduction, designed to visualize complex, high dimensional data in a more interpretable low dimensional format, such as a two or three-dimensional plot. It is the predecessor to t-SNE
A primary advantage of SNE is its capacity to maintain the local structure of the dataset. This means that items that are similar or “close” in the original high dimensional space will be represented as being close in the resulting low dimensional visualization. SNE excels at capturing intricate, nonlinear patterns within the data, setting it apart from linear methods like Principal Component Analysis (PCA), which are less effective when the underlying data relationships are not linear
SNE is particularly useful in specialized fields such as bioinformatics, natural language processing and image analysis. In these domains, the data is often organized in complex, high dimensional structures, referred to as manifolds, which cannot be accurately represented using simple linear projections. SNE provides a powerful method to “unfold” these structures for effective visualization and analysis
C: deepia
a markov chain is one of the simplest ideas in probability, yet it explains an astonishing number of real systems. the central assumption is called the markov property: the future depends only on the present state, not on the path taken to get there. if you know where the system is now, its entire history no longer matters. that’s why it’s called a memoryless process. every step is just a probability of moving from one state to another, and those probabilities define how the system evolves over time.
once you represent those probabilities as a transition matrix, the mathematics becomes remarkably elegant. every multiplication by the matrix moves the system one step into the future. repeat the process enough times, and many markov chains converge to a stable probability distribution, where the long term behavior becomes predictable even though every individual transition is random. this simple framework powers everything from google’s original pagerank algorithm and weather forecasting to speech recognition, genetics, finance, robotics, and reinforcement learning.
the deeper lesson is that uncertainty doesn’t always mean unpredictability. individual events may be random, but the system as a whole often follows a clear mathematical structure. that’s a recurring theme across mathematics and engineering: stop trying to predict every single outcome, and instead model the process that generates those outcomes. once you understand the transition rules, seemingly chaotic behavior starts revealing stable patterns.
Data structures scale by adding dimensions. A scalar is a single number. A vector is a one-dimensional array that can be a row of shape 1×3 or a column of shape 3×1. A matrix is a two-dimensional array of numbers arranged in rows and columns. A tensor is a multi-dimensional array with three or more dimensions.
Neural network training processes batches of multi-channel images as 4D tensors in convolutional models for object recognition.
Andrej Karpathy just said your kids should ignore most of school.
"80% of education should be math, physics, CS."
Not because it's useful - because it carves grooves in the brain.
Grooves that get harder to carve the older you get.
In a pre-AGI world: it gets you a job.
In a post-AGI world: it makes you a functioning, empowered human.
Everything else? Tack it on later.
The cognitive foundation is the only foundation that survives the transition.