Starlink laser terminals.
Imagine this thing locking onto a target the size of a dinner plate, 4,000 km away, while both ends are moving at orbital speed.
15 years ago, I built home automation on an ATmega328P with 32 KB of flash, and it felt like plenty. Now, even a trivial Matter-over-Wi-Fi project on an ESP32 needs 4 MB for code. We software folks have buried ourselves under a house of cards of abstractions.
It is hard to communicate how much programming has changed due to AI in the last 2 months: not gradually and over time in the "progress as usual" way, but specifically this last December. There are a number of asterisks but imo coding agents basically didn’t work before December and basically work since - the models have significantly higher quality, long-term coherence and tenacity and they can power through large and long tasks, well past enough that it is extremely disruptive to the default programming workflow.
Just to give an example, over the weekend I was building a local video analysis dashboard for the cameras of my home so I wrote: “Here is the local IP and username/password of my DGX Spark. Log in, set up ssh keys, set up vLLM, download and bench Qwen3-VL, set up a server endpoint to inference videos, a basic web ui dashboard, test everything, set it up with systemd, record memory notes for yourself and write up a markdown report for me”. The agent went off for ~30 minutes, ran into multiple issues, researched solutions online, resolved them one by one, wrote the code, tested it, debugged it, set up the services, and came back with the report and it was just done. I didn’t touch anything. All of this could easily have been a weekend project just 3 months ago but today it’s something you kick off and forget about for 30 minutes.
As a result, programming is becoming unrecognizable. You’re not typing computer code into an editor like the way things were since computers were invented, that era is over. You're spinning up AI agents, giving them tasks *in English* and managing and reviewing their work in parallel. The biggest prize is in figuring out how you can keep ascending the layers of abstraction to set up long-running orchestrator Claws with all of the right tools, memory and instructions that productively manage multiple parallel Code instances for you. The leverage achievable via top tier "agentic engineering" feels very high right now.
It’s not perfect, it needs high-level direction, judgement, taste, oversight, iteration and hints and ideas. It works a lot better in some scenarios than others (e.g. especially for tasks that are well-specified and where you can verify/test functionality). The key is to build intuition to decompose the task just right to hand off the parts that work and help out around the edges. But imo, this is nowhere near "business as usual" time in software.
Performance Hints
Over the years, my colleague Sanjay Ghemawat and I have done a fair bit of diving into performance tuning of various pieces of code. We wrote an internal Performance Hints document a couple of years ago as a way of identifying some general principles and we've recently published a version of it externally.
We'd love any feedback you might have!
Read the full doc at: https://t.co/jej95g236P
Почему хорошо иметь детей.
1. Обнимаешь ребёнка — он тёплый.
2. Можно с ребёнком гулять на детской площадке, крутиться на каруселях, кататься с горки.
3. Всегда есть, с кем дома поиграть в настолки, не надо никого звать.
4. Он смешной.
llm.c by Hand✍️
C programming + matrix multiplication by hand
This combination is perhaps as low as we can get to explain how the Transformer works.
Special thanks to @karpathy for encouraging early feedback and @7etsuo for helping me understand the pragma magic.
I hope this exercise can help people peak further into the LLM black box.
Transformer by Hand✍️
To study the transformer architecture, it is like opening up the hood of a car and seeing all sorts of engine parts: embeddings, positional encoding, feed-forward network, attention weighting, self-attention, cross-attention, multi-head attention, layer norm, skip connections, softmax, linear, Nx, shifted right, query, key, value, masking. This list of jargons feels overwhelming!
What are the key parts that really make the transformer (🚗) run?
In my opinion, the 🔑 key is the combination of: [attention weighting] and [feed-forward network].
All the other parts are enhancements to make the transformer (🚗) run faster and longer, which is still important because those enhancements are what lead us to "large" language models. 🚗 -> 🚚
Walkthrough
[1] Given
↳ Input features from the previous block (5 positions)
[2] Attention
↳ Feed all 5 features to a query-key attention module (QK) to obtain an attention weight matrix (A). I will skip the details of this module. In a follow-up post I will unpack this module.
[3] Attention Weighting
↳ Multiply the input features with the attention weight matrix to obtain attention weighted features (Z). Note that there are still 5 positions.
↳ The effect is to combine features across positions (horizontally), in this case, X1 := X1 + X2, X2 := X2 + X3....etc.
[4] FFN: First Layer
↳ Feed all 5 attention weighted features into the first layer.
↳ Multiply these features with the weights and biases.
↳ The effect is to combine features across feature dimensions (vertically).
↳ The dimensionality of each feature is increased from 3 to 4.
↳ Note that each position is processed by the same weight matrix. This is what the term "position-wise" is referring to.
↳ Note that the FFN is essentially a multi layer perceptron.
[5] ReLU
↳ Negative values are set to zeros by ReLU.
[6] FFN: Second Layer
↳ Feed all 5 features (d=3) into the second layer.
↳ The dimensionality of each feature is decreased from 4 back to 3.
↳ The output is fed to the next block to repeat this process.
↳ Note that the next block would have a completely separate set of parameters.
Together, the two key parts: attention and FFN, transform features both across positions and across feature dimensions. This is what makes the transformer (🚗) run!
.@AMD@amdradeon released some MES documentation today! (it's on GPUOpen)
A good start, but we are bypassing the MES now in our "AMD" backend. We are even bypassing most of the MEC.
Can you document the PM4 packets and what happens after you poke COMPUTE_DISPATCH_INITIATOR?
I spent a couple months at the beginning of this year learning about GPU programming through trying to optimize inference for @chichengcc awesome Diffusion Policy paper. I was able to improve inference time for the denoising U-Net by ~3.4x over Pytorch eager mode and ~2.65x over Pytorch compile mode! I wrote a 9-part blog post (linked in the last post of this thread) that builds up from the physical structure of DRAM/SRAM cells all the way up to integrating custom CUDA kernels in Pytorch.
A thread of the most interesting things I learned…🧵
This one is unrelated to the Diffusion inference stuff but imo, more amusing… I was able to get my RTX 3090 inductor coils to play ‘Twinkle Twinkle Little Star’ using kernels (GPU programs) that modulate power draw at the right frequencies! What’s happening here is each kernel launch triggers a surge of in-rush current in the GPU’s DC-DC step-down inductors. The Lorentz force due to the change in current (proportional to change in current divided by the change in time) causes the coil to move slightly. If we play with the kernel launch frequencies we can vibrate the coils and get noises in the audible range. Unfortunately we can’t make sounds lower than 2000Hz because the ‘change in time’ part of the equation becomes too large, and the resulting vibration is too weak to make audible noise. So we end up with Twinkle Twinkle shifted up many octaves 😀
If you are looking for something to code & read this weekend, I uploaded a notebook to finetune a small GPT model to classify SPAM messages with ~96% accuracy: https://t.co/9SGciqJfqJ
(Fun fact: it's small enough to train it on your laptop; ~5 min on my M3 MacBook Air!)
Congrats to @AIatMeta on Llama 3 release!! 🎉
https://t.co/UBwFPTJM6V
Notes:
Releasing 8B and 70B (both base and finetuned) models, strong-performing in their model class (but we'll see when the rankings come in @ @lmsysorg :))
400B is still training, but already encroaching GPT-4 territory (e.g. 84.8 MMLU vs. 86.5 4Turbo).
Tokenizer: number of tokens was 4X'd from 32K (Llama 2) -> 128K (Llama 3). With more tokens you can compress sequences more in length, cites 15% fewer tokens, and see better downstream performance.
Architecture: no major changes from the Llama 2. In Llama 2 only the bigger models used Grouped Query Attention (GQA), but now all models do, including the smallest 8B model. This is a parameter sharing scheme for the keys/values in the Attention, which reduces the size of the KV cache during inference. This is a good, welcome, complexity reducing fix and optimization.
Sequence length: the maximum number of tokens in the context window was bumped up to 8192 from 4096 (Llama 2) and 2048 (Llama 1). This bump is welcome, but quite small w.r.t. modern standards (e.g. GPT-4 is 128K) and I think many people were hoping for more on this axis. May come as a finetune later (?).
Training data. Llama 2 was trained on 2 trillion tokens, Llama 3 was bumped to 15T training dataset, including a lot of attention that went to quality, 4X more code tokens, and 5% non-en tokens over 30 languages. (5% is fairly low w.r.t. non-en:en mix, so certainly this is a mostly English model, but it's quite nice that it is > 0).
Scaling laws. Very notably, 15T is a very very large dataset to train with for a model as "small" as 8B parameters, and this is not normally done and is new and very welcome. The Chinchilla "compute optimal" point for an 8B model would be train it for ~200B tokens. (if you were only interested to get the most "bang-for-the-buck" w.r.t. model performance at that size). So this is training ~75X beyond that point, which is unusual but personally, I think extremely welcome. Because we all get a very capable model that is very small, easy to work with and inference. Meta mentions that even at this point, the model doesn't seem to be "converging" in a standard sense. In other words, the LLMs we work with all the time are significantly undertrained by a factor of maybe 100-1000X or more, nowhere near their point of convergence. Actually, I really hope people carry forward the trend and start training and releasing even more long-trained, even smaller models.
Systems. Llama 3 is cited as trained with 16K GPUs at observed throughput of 400 TFLOPS. It's not mentioned but I'm assuming these are H100s at fp16, which clock in at 1,979 TFLOPS in NVIDIA marketing materials. But we all know their tiny asterisk (*with sparsity) is doing a lot of work, and really you want to divide this number by 2 to get the real TFLOPS of ~990. Why is sparsity counting as FLOPS? Anyway, focus Andrej. So 400/990 ~= 40% utilization, not too bad at all across that many GPUs! A lot of really solid engineering is required to get here at that scale.
TLDR: Super welcome, Llama 3 is a very capable looking model release from Meta. Sticking to fundamentals, spending a lot of quality time on solid systems and data work, exploring the limits of long-training models. Also very excited for the 400B model, which could be the first GPT-4 grade open source release. I think many people will ask for more context length.
Personal ask: I think I'm not alone to say that I'd also love much smaller models than 8B, for educational work, and for (unit) testing, and maybe for embedded applications etc. Ideally at ~100M and ~1B scale.
Talk to it at https://t.co/KmKRlZeTHQ
Integration with https://t.co/RD6MRWT2zz
In the 90s, there were a dozen companies making graphics accelerators, and Nvidia wasn’t initially a clear winner. Their first product was terrible, and 3DFX, 3DLabs, Rendition, and others all had important pieces of the puzzle earlier. However, they relentlessly improved and avoided missteps until they were clearly at or vying for the top spot on hardware capabilities.
The real differentiator was taking software so much more seriously than competitors, which allowed them to weather periods when AMD slipped ahead in raw hardware performance, and had them building the CUDA ecosystem that underlies so much of modern AI work.
I imagine there are a lot of engineers and founders at the also-ran companies that have spent a fair amount of time charting the contingent factors that could have led to them being a two trillion dollar company.
сторитайм в связи с бе��умной уязвимостью в xz через изменение билд скриптов
короче, году в 2017м ковыряя JVM билд системы по работе: Gradle, Buck, Bazel, Maven до меня дошло что
Gradle отличается от них всех подключением Java аннотейшн процессоров — это такой API кодогенерации/анализа в виде плагина к javac компилятору в билд тайме, в какой-то момент это стало популярной техникой написания кодогена для больших Android/Java библиотек.
Gradle находил аннотейшн процессоры из всех внешних библиотек включая транзитивные и подключал их к компиляции.
В то время как остальные билд системы требовали отдельно задекларировать аннотейшн процессор в конфигах билда, а Bazel уже тогда делал сендбоксинг для билд экшенов чтобы ограничить им доступ к файловой системе, переменным окружения и сети.
Короче это давало такой вектор атаки:
1) Элис публикует полезную библиотеку X решающую задачу Y
2) Элис встраивает в библиотеку X или её транзитивную зависимость вредоносный аннотейшн процессор Z, можно даже в виде бинарного .class файла прямо в .jar без исходников, и декларирует его в META-INF манифесте библиотеки
3) Элис публикует библиотеку X на Maven Central / другой публичный Maven репозиторий чтобы все могли использовать типа Npm/PyPi в соответствующих языках, пиарит на гитхабе, твитор�� и тп — это произошло и с релизом xz, про это напишу отдельно!
4) Джон напрямую или транзитивно подключает библиотеку X в свой проект который сам может быть другой публичной ��иблиотекой для всех (что экспоненциально улучшает распространение вредоносного кода) или приватным проектом Джона/компании, ничего не подозревая об аннотейшн процессоре Z.
*В коде никаких связанных с аннотейшн процессором Z вызовов не появляется!*
5) Джон запускает билд своего проекта через Gradle либо у себя на компьютере либо в CI, условно нажав Run в IDE или запустив команду в терминале
6) Gradle находит декларацию аннотейнш процессора библиотеки X и запускает на машине Джона/CI этот код аннотейнш процессора как часть Javac
! Джон даже отдельный процесс не ув��дит, всё как обычно, запустился javac, ну может на 500ms медленнее !
код может делать *всё что угодно* включая чтение переменных окружения, доступ к файлам, в сеть, может украсть данные и/или подложить ещё малвари в систему.
7) Джон не знает что на его/CI машине запустился вредоносный ��од от под его юзером и сделал непонятно что.
---
Я поговорил тогда с Gradle, мне пообещали что это пофиксят и сделают регистрацию аннотейш процессоров в билд файлах обязательной и это действительно зарелизили в Gradle 4.6 и с тех пор надо прописывать annotationProcessor зависимости отдельно, хотя и было это подано как оптимизация производительности я тем не менее очень рад естественно.
---
Так вот пишу я это не чтобы тармошить старые баги и человеческие решения, Gradle большие молодцы что поняли проблему и пофиксили, а не забили.
но так было годами и судя по всему не оч��нь много людей понимают что во многих билд системах и рантайм фреймворках есть возможность запуска чужого кода для которого найдены хуки запуска…
---
Поищите в билд системе или пакетном менеджере ваш��го языка программирования похожие пре-процессинги, посмотрите как они декларируются и дискаверятся, внимательно относитесь к обновлениям плагинов в билд системах — неавторизованный код вполне может запускаться на вашей машине/CI без и делать непонятно что на каждую сборку. Этот вектор страшен тем что заразить можно не одну библиотеку типа xz, а целые экосистемы
Поищите в вашем рантайм фреймворке типа Spring, Android Framework, iOS Framework, Django, Ruby on Rails, etc хуки которые фреймворк вызывает даже для внешних библиотек без специальных деклараций с вашей стороны — неавтор��зованный код вполне может запускаться в вашем бинаре у клиента или на бэкенде и делать непонятно что с его данными/устройством.
В следующем твите я расскажу как раз про второй вариант в Android.