Why LLM reread unchanged data?
So I applied video compression to LLM inference: consecutive KV cache values are nearly identical, so I store deltas instead of full values. At the same 4-bit precision, error drops 10,000×.
https://t.co/ct6isFmjRp
p.s. my mi50 cluster