Here's something I keep wondering about: we quantize models to 4 bits and they barely notice. But push toward 2 bits and the quality starts slipping fast. What decides where that line sits?
Chasing this took me back to Claude Shannon's 1959 rate-distortion paper. A theory built for copper wires, possibly deciding today whether a language model fits on your phone.
And one recent result complicated the picture in an interesting way: models trained on more tokens appear harder to compress. Almost as if we're training them out of compressibility.
https://t.co/bFT7Kfea42
Here's something I keep wondering about: we quantize models to 4 bits and they barely notice. But push toward 2 bits and the quality starts slipping fast. What decides where that line sits?
Chasing this took me back to Claude Shannon's 1959 rate-distortion paper. A theory built for copper wires, possibly deciding today whether a language model fits on your phone.
And one recent result complicated the picture in an interesting way: models trained on more tokens appear harder to compress. Almost as if we're training them out of compressibility.
https://t.co/bFT7Kfea42
Tencent has killed fine-tuning and RL with a $18 budget.
Right now, if you want an AI agent to become an expert at a specific, complex real-world task, you have to use Reinforcement Learning.
You let it try, fail, and update its internal parameters over and over again.
This is the exact optimization technique (GRPO) that DeepSeek used to build their massive reasoning models.
But there is a massive problem.
Updating model weights is insanely expensive. It requires massive GPU clusters. And worst of all, when you train a model to be highly specialized at one thing, it often "overfits" and forgets how to be good at everything else.
Tencent killed this bottleneck forever.. by building Training-Free GRPO.
Instead of spending thousands of dollars to permanently alter the AI's brain, they asked a simple question: What if we just distill the experience of learning, and inject it as a memory?
Here is how it works.
They run the AI through the exact same trial-and-error process. But instead of updating the weights, they extract the "semantic advantage"—the actual logic of why one answer was better than another.
They compress this winning logic into a "token prior”, a tiny package of high-quality experiential knowledge.
Then, they just attach that knowledge directly into the API call.
The results are staggering.
Tested on DeepSeek-V3, this method required only a few dozen training samples to turn the AI into a specialized expert in complex math and web searching.
It didn't just compete with models that were actually fine-tuned. It outperformed them.
Zero parameter updates. Zero expensive training runs. Zero base-model amnesia.
@VFSGlobal I submitted my tatkal application for passport renewal through VFS NYC with the missing docs and still got no response. Their systems are very outdated and each customer care person has different reasons to keep my application on hold. VFS has the WORST SERVICE EVER