Caught up with life and work for the past few days... But, instead of CUDA, been learning about Agentic AI. I'll share a post of what I've covered so far, ranging from ReAct loops, context engineering, memory, etc.
Continued with CUDA today. Practiced previous kernels and wrote 2 new kernels:
1. Convolution 2D
2. Softmax 2D
These are still the naive implementations, but the idea so far is simple: assign each GPU thread a calculation/operation that can be executed parallelly.
#CUDA
We overestimate what we can do in a day and underestimate what we can do in 3 months through focused, consistent effort.
That is my goal. To learn about inference and CUDA, every day, bit by bit, by showing up consistently.
Continued with my learning and coded the Causal Attention class. This prevents a token from accessing tokens in the future through a mask. A dropout is also applied, so that overfitting can be prevented. The dropout mask is applied after computing attention weights.
#ML
@pavelsimo Following along! It's been 4 days since I started my CUDA journey. Just wrote a naive GEMM kernel yesterday. Still a long way to go haha. Hopefully I'll get to learn something from your journey!
Should we still learn how to code?
If you want to clear interviews, then absolutely. You need to have a language of your preference, and know the inner workings of it.
At your job, mostly no, since coding agents will write most of the code.
Every mature systems project eventually needs a teaching version.
TinyTorch is a free, open source curriculum where you build a working machine learning framework from scratch, tensors through transformers, using PyTorchโs own API in pure Python.
It requires no GPU, runs on a 4 GB laptop, and covers 20 hands-on modules designed to give developers, students, and engineers a complete mental model of PyTorch internals.
Read the full technical breakdown here: https://t.co/UFeJKNWKmO
@profvjreddi
I also dove into the mathematical proof behind why do we scale attention scores. Yeah, numerical stability before feeding it into softmax is a good reason, but why exactly with sq_root(d_k)?
Turns out, it has everything to do with variance and expectation. Neat stuff!
#ML
Today, I pivoted from CUDA to learn about self-attention. The aim of self-attention is to create context vectors, which is a much richer representation of a token in a given sentence.
@h00recki Hey thanks! Not yet, but I plan to ๐ช. Currently i am writing cpu functions and converting them to naive gpu kernels just so that i get the hang of indexes.
Also its awesome to hear that you are getting into it!
The famous GEMM! I finally wrote my first GEMM cuda kernel today. This is the naive version of it, but proud that I'm making consistent progress in cuda :)
#cuda
@MicrosoftLearn Learning about CUDA and LLM inference lately. Taking 30-60 mins every day, making small, incremental steps. Gotta stack those receipts!
8/
While some attention architectures tried to reduce the KV cache footprint, some tried to limit how much prefix a token can attend to (Sliding window and Sparse attention).
It is all about inference.
Still a lot to learn.
Spent some time today going down the rabbit hole of attention mechanisms.
Started with the familiar self-attention โ MHA, and ended up encountering things like MQA, GQA, Multi-Head Latent Attention (MLA), Kimi Delta Attention, and hybrid attention architectures.
๐งต
7/
The article also covered Sliding Window Attention, Kimi Delta Attention and hybrid attention architectures.
I definitely didn't understand all the details, but I liked seeing the progression:
MHA โ MQA โ GQA โ MLA โ newer hybrid designs.