I remembered a moment before this war…
Whenever the Israeli army invaded my town, Beit Hanoun, my friends and I would run into the streets and throw stones toward the soldiers.
They fired bullets at us, while we had nothing but stones in our hands.
I knew that a stone was nothing compared to a weapon, but because it was my land… I felt that the stone was stronger than bullets.
That is why I never stopped.
We fought with whatever we had, throwing those stones as if we were defending everything we loved. 💔
Nestle Msia. Terkenal pelbagai produk. Always resilient.
Khamis lps Nestle umum 2Q profit hanya RM93mil, paling terok tempoh 8 thn. Revenue RM1.5bn lowest since 2021. Menurut analysts “consumer spending weakened” dan lemah sebab “heightened inflation pressure”.
Saham fell 8% in 2 days. Lowest close since 2018.
These 94 lines of code are everything that is needed to train a neural network. Everything else is just efficiency.
This is my earlier project Micrograd. It implements a scalar-valued auto-grad engine. You start with some numbers at the leafs (usually the input data and the neural network parameters), build up a computational graph with operations like + and * that mix them, and the graph ends with a single value at the very end (the loss). You then go backwards through the graph applying chain rule at each node to calculate the gradients. The gradients tell you how to nudge your parameters to decrease the loss (and hence improve your network).
Sometimes when things get too complicated, I come back to this code and just breathe a little. But ok ok you also do have to know what the computational graph should be (e.g. MLP -> Transformer), what the loss function should be (e.g. autoregressive/diffusion), how to best use the gradients for a parameter update (e.g. SGD -> AdamW) etc etc. But it is the core of what is mostly happening.
The 1986 paper from Rumelhart, Hinton, Williams that popularized and used this algorithm (backpropagation) for training neural nets:
https://t.co/f52IcDNitR
micrograd on Github: https://t.co/GaTd16jRnB
and my (now somewhat old) YouTube video where I very slowly build and explain:
https://t.co/EPGG6kd5Yz
Have you ever wanted to train LLMs in pure C without 245MB of PyTorch and 107MB of cPython? No? Well now you can! With llm.c:
https://t.co/PoGTZIwASL
To start, implements GPT-2 training on CPU/fp32 in only ~1,000 lines of clean code. It compiles and runs instantly, and exactly matches the PyTorch reference implementation.
I chose GPT-2 to start because it is the grand-daddy of LLMs, the first time the LLM stack was put together in a recognizably modern form, and with model weights available.
no matter how chaotic and unstable the orbits of these three gravitational bodies are, their centre of gravity (centroid of the triangle) always remains stationary and fixed in space
Starting to think we should drop the "Large" label in LLM:
*Small vs. large range is all over the place. Bert was the og LLM at 170M, but now seeing people arguing that <100B is small
*Mixture of Experts blur the concept of size, with models being simultaneously 7B, 14B and >50B
You know what the biggest problem with pushing all-things-AI is? Wrong direction.
I want AI to do my laundry and dishes so that I can do art and writing, not for AI to do my art and writing so that I can do my laundry and dishes.