@thsottiaux Codex was waiting the running CI runner by loop checking, several percentage week buffer is ated. what's wrong? why don't use event drived monitor?
It is International Cat Day today
Celebrated every year on August 8, the day celebrates cats and raises awareness about their welfare, care and protection
We accidentally built self-improving AI systems. This paper from University of Oxford proves it.
Most people assume model improvements come from bigger architectures or carefully designed reinforcement learning pipelines.
This work shows something more subtle and more unsettling.
If you deploy a model, let users interact with it, filter out the failures, and fine-tune only on the successful traces, the model starts improving its planning abilities on its own.
No explicit rewards, hand-crafted curriculum and no external planner.
Just iteration.
The authors call this iterative deployment, and they test it in controlled planning environments like Blocksworld, Rovers, and Sokoban.
The setup is simple:
1. Deploy an LLM on planning tasks
2. Keep only the plans that actually work
3. Fine-tune the next version on those valid traces
Repeat
After just five generations, planning performance more than doubles across all domains. In some cases it improves by 4 to 5x. Even more interesting, later generations discover much longer plans than the base model, showing real out-of-distribution generalization, not just formatting tricks or prompt compliance.
Here is the key insight.
The paper proves that this process is mathematically equivalent to reinforcement learning with a binary reward signal.
But the reward function is never written down.
It is implicitly defined by user behavior and curation.
Supervised fine-tuning on “only the good outputs” turns out to be REINFORCE in disguise.
That has two big implications.
First, iterative deployment is a powerful alternative to explicit RL for improving reasoning and planning. It works even when rewards are hard to define, as long as you can validate outcomes.
Second, and more worrying, the reward function shaping future models is opaque. User preferences, platform incentives, and validation biases silently become training signals. Over time, these signals can override or conflict with alignment objectives set during pretraining.
In other words, models do not just learn during training.
They keep learning after release.
And they learn whatever the world rewards them for.
This paper reframes deployment itself as a training loop. Once you see it, you cannot unsee it.
Read the full paper: https://t.co/6NfyJ7v91s