How to read the Attention formula from the Attention is all you need paper
Attention(Q, K, V) = softmax( (Q Kᵀ) / √dₖ ) V
This equation shows how every position in a sequence gathers information from every other position.
The diagram walks through the full path from raw inputs X to the final combined context vectors.
"Mathematical Theory of Deep Learning" is another excellent free resource if you are interested in the mathematical analysis of deep neural networks.
The book starts with the formal definition of feedforward neural networks and then covers universal approximation, splines, ReLU networks, high-dimensional approximation and interpolation. It also has detailed chapters on neural network training, gradient descent, stochastic gradient descent and backpropagation. The later chapters are about the geometry of neural network spaces with sections on overparameterization, double descent, robustness and adversarial examples.
It is a solid and very rich resource, with many useful mathematical abstractions for understanding the mechanisms behind deep learning. I think it is especially useful if you want to go beyond a practical use of neural networks and have a more rigorous view of their mathematical foundations. The latest version was updated in January 2026 and is freely available as a PDF.
https://t.co/l8B176FeNl
Elon Musk:
"Don't pursue money.
Make useful products, and money will come as a consequence."
A simple philosophy that has shaped some of the world's most ambitious companies.
I have fallen away from serving Your lotus feet.
O Mother, O Saviour of all, O Auspicious One (Śive)!
Please forgive this.
A bad son may sometimes be born,
but a bad mother can never exist.
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry.
Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential.
OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose.
@rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think!
Try it out: https://t.co/P0mGnI1o31 (requires your own API key)
Source code: https://t.co/NYCiTD6hSq
Erling Haaland doesn’t sleep like a normal person.
He treats sleep like a performance drug.
Just 6 body signals that tell his nervous system it's safe to sleep:
1. Wear blue-blocking glasses 3 hrs before bedtime.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb