Took a while, but here it is, a mega write-up on language models for text classification and, of course, Jev.
It's basically a visual guide to RNNs, CNNs, transformers, and calibration, with hands-on experiments on accuracy and efficiency.
Why KV cache stores K and V vectors but never Q?
(a popular technical LLM interview question)
LLMs are autoregressive so each token is predicted from every token before it, one at a time.
This autoregressive nature has a direct consequence inside the model.
A forward pass over <n> tokens produces <n> hidden states, but only the last one is projected to logits and is required to generate the next token.
So to understand why KV cache just stores K and V vector, we must back track to see how exactly is the last hidden state produced.
Let's walk through this with a 10-token prompt.
1) Prefill:
All 10 tokens go through the model in one forward pass, in parallel (with causal masking), since the whole prompt is already known.
At every layer, each of the 10 positions produces a query, a key and a value vector, and attention at each position runs against all positions up to it.
This pass is compute-heavy, and it's why the first token takes noticeably longer than the ones after it. TTFT is mostly prefill.
2) The first output token:
To generate the 11th token, only the 10th token's hidden state is needed. So this is projected from the hidden-dim to vocab-dim to generate logits over vocab.
These logits then go through softmax and sampling to generate token 11.
3) Back-track the hidden state:
The last hidden state is the last row of the feedforward block's output. The feedforward block is position-wise (it's applied to each row independently) so that row comes from the last row of the attention output before it.
So now we need to see how the last row of attention is computed.
4) Attention matrix:
QKᵀ for a 10-token prompt will give a 10 × 10 matrix.
Row <i> will have the dot product of query <i> with every key.
Row 10 is therefore Q₁₀·K₁, Q₁₀·K₂, all the way to Q₁₀·K₁₀.
Notice that only Q₁₀ appears in it. Q₁ through Q₉ only belong to their corresponding rows 1-9, and those rows' hidden states we already discarded because they were never needed.
The last row of attention goes through softmax and multiplies the full stack of value vectors, V₁ through V₁₀, to give the last row of the attention output.
So the last hidden state depends on exactly three things: Q₁₀, every key, and every value.
5) Generating token 12:
Token 11 is appended, and this time, we need row 11's hidden state to generate token 12.
Mathematically, attention operation turns out to be Q₁₁ against K₁ through K₁₁, then multiplied by V₁ through V₁₁.
K₁ through K₁₁ and V₁ through V₁₁ are bit-for-bit what prefill + first token produced since under causal masking, a token's key and value depend on that token and the ones before it, never on anything after, so appending token 11 cannot change anything at position 3.
6) The cache state:
Overall, this implies that you just need to retain the keys and values at each decoding step, and compute only the new position's Q, K and V.
Each decode step requires one query vector, which is never used again, so they are never cached across the decoding process.
The visual below explains the entire process.
That said, KV cache is only one of four separate caching layers in an LLM stack.
The other three are prefix caching on the server, prompt caching billed by a provider, and a semantic cache that skips the model entirely.
I wrote a full breakdown of all four caches in LLM serving that you should know as an AI engineer, with code for each.
Read it below.
A few algorithms I regularly force myself to review so I don’t forget how they work:
• Kadane’s → Maximum subarray sum
• Rabin-Karp → Substring search with hashing
• Topological Sort → Ordering a DAG
• Prim’s → Minimum Spanning Tree
• Kruskal’s → MST with Union-Find
• Dijkstra’s → Shortest path (no negative weights)
• Bellman-Ford → Shortest path (with negative weights)
• Tarjan’s → Strongly Connected Components
• Backtracking → Sudoku Solver
These are the ones I keep coming back to years later.
Which algorithms do you still periodically revisit?
For many who watched this video, it was a warning from the future.
Over 8 million views across YouTube and Instagram.
One of the most raw encounters I’ve ever had.
👇🏼
new post on harness engineering for AI self-improvement: https://t.co/ZYvGfVs61k
It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple.
Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
CBSE படித்து NEET, JEE-ல அதிக மார்க் எடுத்து அரசு கல்லூரி, IIT, NIT போகும் 5% மாணவர்களைத் தவிர மீதி அனைவருக்கும் CBSE படிப்பு பயனில்லை. TN இன்ஜினியரிங் கல்லூரிகளுக்கு CBSE மார்க் போதாததால், பெரும்பாலானோர் மேனேஜ்மெண்ட் கோட்டாவுக்கே போக வேண்டிய நிலை. 😭😭
A senior Anthropic engineer just dropped 11-page PDF on "Loop Engineering" for agentic systems.
The shift: you stop prompting the agent. You build the system that prompts it instead.
Schedule → Discover → Build → Verify → Repeat
Every loop runs one turn, five moves:
• Discovery: it finds its own work - failing CI, open issues, recent commits - instead of being handed a list.
• Handoff: each task gets an isolated git worktree so parallel agents don't collide.
• Verification: a second agent, told to assume the code is broken, reviews the first. The "thing that can say no."
• Persistence: results get written to disk, never left in a context window that gets flushed.
• Scheduling: an automation wakes it on a timer. That's what makes it a loop.
The key insight: an agent grading its own work always praises it.
This 11-page PDF changed how I'm building agentic systems today.
Read it now, then explore the article below.
Suppose you have interviews scheduled for these 70LPA-1Cr+ CTC roles:
Google L5 /
Meta E5 /
Uber Staff /
Amazon L6 /
Salesforce SMTS /
Go through these 27 core system design concepts and problems.
How many can you reason through in a 1-hour round if the interviewer injects one of them into your design or asks as a follow-up? How much clarity do you have?
Beginner:
- The thundering herd problem
- Cache stampede
- N+1 query problem
- Hot partition / hot key
- Single point of failure
- Retry storm
- Backpressure
- Duplicate requests / idempotency gap
- Stale cache / read-after-write inconsistency
Intermediate:
- Distributed rate limiting
- Leader election
- Distributed locking and lease expiry
- Quorum reads vs quorum writes
- Fan-out on write vs fan-out on read
- Out-of-order event processing
- Dead letter queues and poison messages
- Zero-downtime schema migration
- Circuit breaker and cascading failure control
Advanced:
- Split a monolith safely
- Multi-region failover
- Active-active conflict resolution
- Change Data Capture vs dual writes
- Search index freshness vs ranking quality
- Rebalancing shards under skewed traffic
- Noisy neighbor problem in multi-tenant systems
- Watermarks / late-arriving events in stream processing
- Exactly-once processing vs practical deduplication
Most candidates prepare for:
- “Design Uber.”
- “Design Twitter.”
- “Design Dropbox.”
These generic prompts are good for beginners, but at senior levels, depth matters. Strong candidates prepare for the even minute problems that can cause disasters at scale.
One test you should not ignore is hs-CRP. It often predicts heart disease better than LDL alone. High inflammation = unstable plaques. Test it + act on it! #MetabolicHealth#HeartHealth
Lipid Profile vs ApoB & ApoA1: Which Better Predicts Future Heart Disease Risk?
Most people are familiar with cholesterol testing. But increasingly, we are paying attention to two proteins:
🔸Apolipoprotein B (ApoB)
🔸Apolipoprotein A1 (ApoA1)
What are they and should you get them tested?
A thread.
1/n
Last week I introduced my multi-agent workflow using the Hermes Kanban, in today's video I showed how to actually build one from my open source templates! This workflow is for updating and maintaining knowledge bases, which I think a lot of people could use. Check it out!
https://t.co/ZMgacRGQvm