@mattpocockuk Diff size is a lagging signal. A small change can still invent a term the agent treats as settled, and the next session inherits that assumption. The cheaper check is whether the change needs new shared language (Human & Agent), not how many lines it touches.
Agents don’t learn the way people do. The weights stay the same. What changes is the notes.
It tries, something fails, it writes a short lesson (“validated before conversion, large amounts slipped through”), stores that, and reads it on the next run. That’s Reflexion-style memory, not training. Wipe the store and the lesson is gone. But what if we can store it?
Useful for not repeating the same error. Not the same as the skill sticking in your head because you did the work.
The equilibrium test is the right filter. Most “AI productivity” gains are just temporary arbitrage on lagging adoption. Once everyone has the same tools, the remaining edge is usually in proprietary context, the quality of the problem you point the models at, and how tightly you close the feedback loop. Inbox triage and comment generation fail that test; designing experiments that were previously too expensive, or continuously refining a system with private data that competitors can’t see, still passes it.
The weirdest thing about AI agents:
The more autonomy you give them, the more important your stop conditions become.
A human knows when to give up and ask for help.
Teaching an agent that behavior might be harder than teaching it to complete the task.
The real fix for agent drift isn't manual renaming—it's continuous evaluation. If you treat prompts and modular skills as code, a robust regression test suite (using frameworks like Promptfoo or Braintrust) catches the drift in CI/CD before the files ever need to live or change in production.
@mattpocockuk Love this. Half the battle with long-running agents is stopping them from hallucinating answers to their own backlog before you can actually step in and unblock them.
People say they want honesty.
Until the honest answer isn't the answer they wanted.
Maybe emotional maturity is learning to hear something uncomfortable without immediately turning it into a fight.
One thing I've learned about people:
Sometimes they don't need advice.
They need someone to listen without immediately trying to fix them.
Maybe that's something we'll have to teach AI too.
Not every problem is asking for a solution.