Tweet in Thai/English | 24/7 Depressed and crying | | I’ve been waiting on death with a smile on my face | ไม่มีทางลัดที่จะให้ต้นไม้เติบโต | book lover📚|
bagi yang mau pdf nya, ini yaaa https://t.co/YIgQcjIxph
itu buku daging bgt, ngejelasin identitas matematika secara ringkas terus langsung dikasi latihan soal 🤩
daftar isi bukunya ada di tiga gambar2 ini yakk
5 LLM quantization techniques, clearly explained:
(bookmark this)
A 70B model in FP16 needs 140GB for weights alone. At 4-bit, that drops to 35GB, which fits on one card.
But naive rounding fails on large models. Roughly 0.1% of hidden dimensions carry values up to 20x larger than anything else in the tensor, and they wreck the quantization grid for everything else.
Each of these 5 methods handles those outliers at a different point:
1. RTN: ignores them. Rounds every weight to the nearest grid level with no calibration data. Cheapest option, weakest at low bit widths.
2. GPTQ: repairs after rounding. Quantizes a layer column by column and adjusts the remaining weights to absorb the error before moving on.
3. AWQ: protects before rounding. Finds the ~1% of weight channels that matter most and scales them up so they survive quantization. Everything still ends up in plain INT4.
4. LLM. int8(): isolates at inference. Outlier dimensions run in FP16, the other 99.9% run in INT8, and the results are merged.
5. QAT: solves it during training. The model is fine-tuned with rounding baked into every forward pass, so it adapts to the damage before quantization is actually applied.
All five produce the same artifact, a model at a fraction of its trained precision. They differ only in where the outlier problem gets addressed.
The visual below nicely summarises these techniques.
As further reading, the article below is a first-principles guide to LLM inference that walks through everything between your prompt and the streamed response, covering tokenization, embeddings, attention, the prefill and decode split, KV caching, and quantization.
It will give you a complete mental model of how inference actually works under the hood.
Read it below.
Andrej Karpathy called it a compiler, not a storage system. this diagram shows the architecture he described.
most people who build a second brain treat it like a filing cabinet. they drop notes in. they search when needed. the library gets bigger over time but never smarter.
a compiled wiki works differently. raw holds the source materials. unstructured. immutable. ground truth. wiki is where the model converts everything in raw into structured, linked, evergreen knowledge. the human reads it. the model writes it. output is where finished work lands. built from compiled knowledge, not from memory.
at the center: CLAUDE.md. identity, preferences, goals, project context. the model reads it before every session. automatically. you never explain yourself again.
the loop closes through update. every new source gets ingested, compiled, linked, and integrated. the system improves with every cycle.
five automations run the whole thing. ingest captures and extracts. write retrieves and drafts. manage links decisions to context. review summarizes and reflects. maintain prunes and improves.
the system improves with every cycle. automatically. retrieval answers questions. compilation builds understanding.
the article above. build the compiler. not the cabinet.
a fully open-source self-improving harness.
Prime Intellect just shipped Prime Agent, which turns a frontier model into a Recursive Language Model.
context becomes a variable the model programs over, and sub-agents become ordinary function calls.
let's understand what all of this means:
almost every agent you use today works inside fixed harness. someone wrote the system prompt, picked the tool schemas, decided when history gets compacted, and shipped it.
the model spends part of every session working around that scaffolding instead of working on the task. when the same failure repeats three times in a long run, you are the one who fixes it, after the run is over.
Prime Agent removes that ceiling with two pieces.
the first is the RLM design. a persistent IPython kernel is the model's only tool, so long inputs never have to enter the prompt at all. the model greps, partitions, and spawns child calls over the data instead of reading all of it back in.
the second is the Continual Harness. four kinds of state sit outside the conversation and stay writable:
→ Prompt: supplemental instructions the agent appends when it learns something the base prompt never told it.
→ Memory: findings from earlier turns that would otherwise die with the context window.
→ Skills: recurring workflows packaged as importable Python, so the next run imports the procedure instead of rediscovering it.
→ Sub-agents: specs for the children it spawns, tuned once and reused across parallel, background, and long-lived runs.
the /refine command reads the current trajectory, applies the smallest edit it can justify, and records what triggered it. the base system prompt stays immutable, and any update can be rolled back by ID.
with Opus 5 driving it, Prime Agent reports 95.5% on ARC-AGI-3, just past the reported human expert baseline of 95.4%. when that benchmark launched, frontier models were scoring under one percent. nobody trained a new model to close that gap.
worth noting that in Factorio runs the same loop found and then optimized scoring exploits, which is roughly what you would expect once an agent can edit its own instructions.
with Opus 5 driving it, Prime Agent reports 95.5% on ARC-AGI-3, just past the reported human expert baseline of 95.4%. when that benchmark launched, frontier models were scoring under one percent. nobody trained a new model to close that gap.
the scaffold stopped being something you configure once and became something the run improves as it goes.
Prime agent is MIT licensed, single command install, works with open and closed models.
check it out on GitHub: https://t.co/h4QSrJNJQE
since we are talking about RLMs, i wrote a detailed article on how Recursive Language Models work.
the article is quoted below.
Harvard, Andrew Ng, and Karpathy will teach you AI engineering for free.
Almost all of it is free, and the order matters as much as the resources.
Most people just do it in the wrong order (here's the correct order):
This paper is f*cking insane
paper hits the core flaw in AI memory engineering: agents pile everything into one flat graph, then drown in junk and rewrite the same facts over and over
The result: HiGram, a hierarchical graph memory that finds the exact evidence path a question needs, then rewrites only that bounded region
The crazy part is how surgical it gets. it edits the affected subgraph and its dependencies together, instead of patching one node and breaking the rest
It's coarse to fine: high-level nodes abstract the topics, MemoryUnits hold the fine facts, so retrieval never traverses the whole graph
Most graph-memory systems update units independently, so one change forces endless rewrites and huge token cost
HiGram beats the baselines on long-term QA and conflict-aware memory: better answers, sharper evidence selection, lower token cost
Read the paper + article below. bookmark it
How to read the Attention formula from the Attention is all you need paper
Attention(Q, K, V) = softmax( (Q Kᵀ) / √dₖ ) V
This equation shows how every position in a sequence gathers information from every other position.
The diagram walks through the full path from raw inputs X to the final combined context vectors.
NVIDIA researchers did it again!
They found a way to make KV cache transferable between models.
The target model skips prefill entirely, and the conversion runs 2.7 to 25x faster than processing the context again.
Let's understand why this is so important today.
LLM APIs are stateless, so every turn sends the entire conversation back to the model. The model reads all of it again before writing a single new token, and all of it is billed as input.
Prompt caching allows Anthropic and other providers to hold the KV cache for a stable prefix and bill a hit at roughly 10% of the base input rate, because the compute was already done once.
The 90% reduction is one of the largest lever in LLM serving, which is why so much production work goes into keeping prefixes byte-stable.
But the cache only works on the model that produced it. Keys and values are produced from that model's weights, so no other model can read them.
In pratice, the constraint shows up in LLM routing. If the traffic is shifted to a different model for cost/capability reasons, the accumulated KV cache becomes invalid.
As a result, the accumulated context has to be processed from scratch, and it's billed at full rate.
NVIDIA's recent paper treats this as a representation problem.
Prefill's only output is the KV cache, so to move KV between models, we need to convert one model's cache into the format the other expects.
They first checked whether the conversion has any structure worth exploiting.
They found that moving from Qwen3 14B to 32B, a plain linear regression from a single source layer reconstructed 56% of the variance in the target model's keys.
The two models obviously may have different layer counts, so there is no natural one-to-one pairing between them.
For each target layer they rank every source layer by how well it predicts that layer, then feed the top eight in together, which takes the reconstruction to 79%.
The mapper itself has three parts:
> Each target layer and head gets its own independent linear map, solved in one closed-form step rather than by gradient descent.
> The cross-layer selection described above is the second part, and their ablation shows it carries the most weight of the three.
> Keys also carry a position-dependent rotation from RoPE. They strip that rotation, fit the map in position-free space, then re-apply the target model's rotation at inference.
Across six pairs from Qwen3, Llama 3.1 and Ministral 3, four retain 73 to 98% of the receiving model's standalone accuracy, and the conversion runs 3-25x faster than processing the context again.
Prior work on cross-model KV reuse exists, but it either trains a neural adapter per pair or requires both models to be architecturally identical.
This is probably the first version that is closed-form and training-free, so a lot of it is still open research.
Every pair tested belongs to one family, so it works on Qwen to Qwen and Llama to Llama.
Cross-family transfer is listed as future work.
All six pairs mentioned above also happen to share KV head count and per-head dimension across scales. Mismatched head configurations are currently untested.
The researchers scoped this to dense full-attention only, so sliding-window and attention-recurrent hybrids still need work.
Here's the paper: https://t.co/tMUGhijFbc
Plenty of work is yet to be done. Still, the constraint being solved is genuine.
Every model swap currently invalidates the full KV that was already paid for, and this is the first result showing that work might be recoverable without training anything extra.
That said, all of this only matters because of what the KV cache is doing in the first place.
I wrote a first-principles breakdown of it, covering why the model stores keys and values at all, why the cache grows with every token, and what generation speed looks like with and without it.
Read it below.
Just saw that the LLMs-from-scratch repository passed 100,000 stars on GitHub!
This is super cool and motivating. I am really happy to see that this open-source repo has helped so many people.
Thanks also to everyone who shared ideas and opened PRs with improvements!
Of course, I plan to keep adding new material, including new attention variants and architectures (while bigger projects like RL and Reasoning From Scratch live in their separate repositories).
I am also currently working on a larger applied custom “small” LLM project. It has been keeping me super busy this month, but I will share more on that soon.
If you are new to it, some of the highlights include
1. Of course, the complete code path from tokenization and attention to pretraining, classification, and instruction fine-tuning, etc. All of it FROM SCRATCH, of course! (RL lives in a companion repo.)
2. From-scratch implementations of Llama, Qwen, Gemma, and Olmo (smaller variants that run locally and can be plugged into the training scripts).
3. From-scratch implementations of attention alternatives and other architecture components, such as GQA, MLA, sliding-window attention, Gated DeltaNet, DeepSeek Sparse Attention, cross-layer KV sharing, and mixture-of-experts
4. Materials on KV caching, training performance, memory-efficient weight loading, DPO, evaluation, and LoRA
So, if you don’t have any weekend plans yet, happy tinkering!
As an AI Infrastructure Engineer.
You can learn:
- GPU/VRAM fundamentals, quantization & batching
- vLLM/TensorRT-LLM/inference optimization
- KV caching, speculative decoding & token throughput
- Distributed training basics (DDP/FSDP/DeepSpeed)
- Model serving & autoscaling
- Vector DB retrieval pipelines
- Prompt caching & cost optimization
- Observability for LLM apps
This is the why production AI teams actually care about.