Great Paper from NVIDIA.
We’ve spent the last couple of years treating RAG and long context as two completely separate engineering headaches.
One is an external pipeline of bi-encoders and rerankers, the other is an expensive brute-force attention problem.
This paper makes a super compelling case that this split is mostly artificial.
Instead of bolting on external retrievers, they tap straight into the frozen LLM’s own internal representations to select evidence across both regimes.
The wild part is that the backbone stays completely frozen.
They only add around 500k trainable parameters.
On a massive 3B-token Wikipedia index, this tiny tweak beats traditional dense and hybrid retriever-reranker stacks across the board (HotpotQA recall jumps from 49% to 73% and 2Wiki leaps from 31% to 60%).
But it also solves the needle-in-a-haystack problem from the other direction.
On long-context benchmarks, using that same internal mechanism to aggressively prune distractors before generation pushes NoLiMA accuracy at 128k context from basically flatlining at 1% up to nearly 25%.
Because it strips out the fluff early, you get a massive drop in FLOPs and time-to-first-token past 32k tokens instead of watching your latency explode.
Great work by Edan Kinderman and the team.
Read the full paper here: https://t.co/tNREGwtKZP
China quietly solved the biggest problem in RL training.
Kimi published a paper that makes LLM reinforcement learning 2× faster.
Right now, every major lab uses Reinforcement Learning to make their models smarter.
But there’s a massive structural flaw in how it works.
During training, the AI is asked to generate multiple different answers to the same prompt so it can figure out which one is best.
Here is the problem: some answers are short. Some answers are long.
Because the system is synchronous, the entire multi-million dollar GPU cluster has to sit and wait for the single longest answer to finish generating before it can update the model.
It’s like keeping the entire class inside until the slowest kid finishes the test.
The GPUs sit idle. Throughput dies.
Kimi researchers dropped a paper introducing "Seer," a system that completely eliminates this bottleneck.
Instead of waiting, Seer uses "Online Context Learning."
It dynamically splits the generation process into movable chunks and learns from the other responses to the same prompt on the fly.
It stops the GPUs from waiting.
The results are staggering.
Seer achieves up to a 97% improvement in training throughput compared to state-of-the-art baselines.
It reduces that dreaded long-tail waiting time by 93%.
No massive architecture overhaul. Just a radically smarter way to manage the flow of data.
CacheFuse is an early attempt in cross-model KV cache reuse that demonstrates that linearly interpolating the KV caches of two models can outperform a standalone model’s cache alone. I’ll be presenting it tomorrow at #COLM Efficient Reasoning Workshop - stop by!
Samsung open-sourced a method that shrinks 13B parameter LLM into less than 1 GB.
It's called "LittleBit"
Instead of storing AI weights as standard numbers, they used latent factorization to crush them down to extreme sub-1-bit levels.
In some configurations, they hit 0.1 bits per weight.
But here is where the architecture gets crazy.
When you compress an AI this much, you can stop doing math.
Instead of forcing the hardware to do heavy floating-point multiplication, LittleBit replaces the core computation with a bitwise XOR operation.
It swaps complex matrix math for basic sign flips.
The results rewrite the rules of model deployment:
• Unlocks a massive 11.6x inference speedup relative to standard FP16 models.
• Radically reduces memory footprint and loading bandwidth.
• Maintains robustness in extreme sub-0.5 bit regimes where previous compression methods catastrophically fail.
This is not a clever optimization.
It is the blueprint for running massive, state-of-the-art AI locally on cheap, resource-constrained devices.
The largest agentic LLM Inference dataset has arrived!
206B tokens · 12,002 sessions · 1,186,582 LLM requests · 1,213,347 tool calls
The traces include tool names, arguments, durations, and more.
Explore the dataset on @huggingface https://t.co/12hrLuupVF
Paper coming soon! Stay tuned for more at https://t.co/8NnRV9xiN5.