KV Cache is one reason LLM inference can be much faster than repeatedly recomputing the entire context.
The model stores previously calculated:
Key + Value states
and reuses them for subsequent tokens.
Less redundant computation.
Much faster generation.
The trade-off?
KV cache consumes memory.
Prompting changes the input.
Fine-tuning changes the model.
One teaches the model what to do for a single request.
The other changes its behavior across every future request.
That's why fine-tuning is much more expensive and much more powerful.
NVIDIA researchers built a new transformer variant.
One small change to the layers made:
- decoding 1.7x faster
- long-reasoning accuracy up 6.5 points
In a typical transformer architecture, every attention layer computes Q, K, and V.
NVIDIA's tweak adds a fourth projection, which predicts what the next layer will need.
To understand why they did this, let's first see what happens in a Transformer architecture during inference right now.
Sparse attention was an attempt to handle long-context inference. Instead of attending to every cached token, modern designs score the KV cache in blocks, keep the top-k, and attend only to those.
That cuts attention compute and bandwidth, but this still leaves us with two problems.
> First, the KV cache still grows with every generated token.
At 100K+ context, it no longer fits in GPU memory and gets offloaded to CPU RAM.
Now every layer must first copy its selected KV blocks from CPU memory back to the GPU. That copy is slow, the GPU sits idle while it waits, and the stall repeats at every layer of every decode step.
> Second, the selection step itself is not free.
Standard selectors score every candidate block with every query head in a GQA group (grouped-query attention, where several query heads share one KV head), then softmax each head's scores and sum them across the group.
During decode, the sparse attention itself is cheap because there is only one query token.
But the expensive part is deciding which blocks to attend to, and that cost keeps growing with context length.
Both problems trace back to the same design in today's sparse attention methods, i.e., the attention query drives the block selection.
Selection needs the query vector Q, and Q only exists once its layer is already running. By then, it's too late to fetch anything early.
The query also drags its multi-head layout into selection, so all that scoring computation runs just to make one top-k decision.
A recent paper from NVIDIA and MIT called SparDA breaks this coupling with one architectural change.
Each layer now emits four projections instead of three:
↳ Q, K, V, and a Forecast.
The Forecast from layer L predicts which KV blocks layer L+1 will need.
Layer L+1's own query performs the sparse attention over those selected blocks.
This one change fixes both problems.
Since the next layer's block set is known while the current layer is still computing, the runtime fetches those blocks from CPU memory on a separate CUDA stream.
The copy overlaps with the current layer's compute, so the GPU no longer waits for it.
And since the Forecast is separate from the attention query, it doesn't need one score per query head.
SparDA uses one Forecast head per GQA group, which removes the per-query-head scoring loop and skips the softmax step entirely.
DeepSeek did something similar in DSA, where a small indexer picks important tokens instead of the query doing it.
SparDA applies the same idea to blocks and adds the prefetch angle that DSA doesn't touch.
The cost of the change is small.
The Forecast adds just 33.5M parameters on an 8B model (0.41%), and only those projections are trained, using a KL loss that matches the original selector's block distribution.
On MiniCPM4.1-8B and NOSA-8B, accuracy matches or beats the sparse baseline, with NOSA-8B gaining +6.5 on long reasoning.
Prefill runs up to 1.25x faster and decode up to 1.7x faster than the sparse offload baseline.
There's one more benefit.
Because prefetch hides the offload cost, most of the KV cache can live in CPU RAM, and the freed GPU memory fits much bigger batches, pushing decode throughput up to 5.3x over the non-offload sparse baseline.
That said, this lookahead will only pay off during decode with CPU offload. During prefill, all keys already live on the GPU, so the gain there comes purely from the cheaper selection.
Here's the paper: https://t.co/wnMG6iRcv9
I wrote a first-principles breakdown of how the KV cache works. It walks through why the model stores keys and values at all, why the cache grows with every token, and a comparison of LLM generation speed with and without KV caching.
Read it below.
Everyone talks about embeddings and vector databases in RAG.
But here's something underrated:
Bad chunking → bad retrieval → bad answers.
Imagine your documentation contains:
Authentication → JWT → Token Expiry → Refresh Token
If you put the entire document into one huge chunk, a query like: "How long does the JWT last?" may retrieve lots of irrelevant context.
Instead, chunk around meaningful sections:
→ Authentication
→ JWT
→ Token Expiry
→ Refresh Token
Now retrieval has a much better chance of finding exactly what matters.
RAG isn't just:
Documents → Embeddings → Vector DB → LLM
It's: Documents → Smart Chunking → Embeddings → Retrieval → Context → LLM
The quality of your RAG system starts before the first embedding is created
Everyone says "Use RAG"
But here's the actual problem it solves.
LLMs don't "look up" answers.
They generate the next token based on patterns learned during training.
That means:
• No access to your latest documents
• No company-specific knowledge
• No guarantee of factual accuracy
RAG changes the flow:
User Query
→ Embed Query
→ Vector Search
→ Retrieve Top-K Documents
→ Add Context
→ LLM Generates Answer
The LLM isn't becoming smarter. It's becoming better informed.
16 AI-designed viruses that do not exist in nature formed viable phages, and a mix of them killed E. coli strains resistant to the natural PhiX174 virus. @sciencemagazine https://t.co/HiYtoOrbOl
0.7 volts from concentration batteries—more than 10 times the assumed limit—shows how engineering ion surroundings can expand electrolyte design without changing electrode materials. @cornell@NatureComms@ScienceAdvances https://t.co/NZOlCkXwBa
Up to 100 meteors an hour may light up the Perseids’ peak, with a new moon setting up darker skies; the brightest fireballs are best seen after midnight and before dawn. https://t.co/zeWicTC0RV
Being around good people can create a positive and nurturing environment, resulting in happiness and fulfillment.
When we are surrounded by individuals who are kind, compassionate, and empathetic, we are more likely to feel supported, loved, and connected.
Day 1 of taking AI seriously.
I'm documenting everything I learn:
• LLMs
• RAG
• MCP
• AI Agents
• Open Source Models
If you're on the same journey, let's learn together.
75% less platinum, nearly the same fuel-cell performance: a compressed copper-platinum catalyst reached 855 millivolts, virtually matching pure platinum’s 856, though long-term stability still needs testing. https://t.co/bbBnVWBg3f