I have a theory I keep failing to break: every kind of human work in history — farming to flight checklists to AI agents — is the same loop.
Convert an uncharted problem into a corridor. Then walk it, repeatedly, with new content.
Name the work that doesn't fit. 🧵
A Terminal-Bench 3.0 task: 59 public runs, 11 model/agent configs, zero passes.
My method got Codex through it. 19/19 official checks, 5,107s of 5,400s.
Harness, verifier, task — untouched.
4th attempt though, method revised in between, no control arm. Data + method open:
https://t.co/ea2voXJJqf
@thsottiaux Built-in drift detection and prevention for long-running coding agents. We model reliability through three probabilistic variables—valid position, valid direction, and valid entrance—and constrain each at runtime. The theory:
https://t.co/v8xHrVA05v
@AndrewYNg Statistical methods to measure, steer, and govern AI are the missing reliability layer for coding agents. Our model frames long-horizon drift detection and prevention through 3 probabilistic variables: valid position × valid direction × valid entrance.
https://t.co/v8xHrVA05v
@ajambrosino MCP protocol
ATM machine
PDF format
PIN number
AI agent drift
The last one isn’t redundant—it’s still missing reliable detection and prevention.
Our model:
P(reliable work) = P(valid position) × P(valid direction) × P(valid entrance)
https://t.co/v8xHrVA05v
Codex is excellent at execution: edit, test, commit, push. What’s missing is drift detection and prevention for long-running work.
Our probabilistic model defines reliability as:
P(reliable work) = P(valid position) × P(valid direction) × P(valid entrance)
Systems can detect and reduce drift by constraining these three variables.
https://t.co/35H8oV3ksY
This lands hard. Graphs as the structurization layer is exactly the “charting” step I’ve been seeing in production.
Once dependencies, legal moves and acceptance criteria are made explicit (the corridor), agents stop hallucinating and start walking reliably. Without it everything stays uncharted and expensive.
I had an agent rewrite its own acceptance rule and run for 15 days before I noticed — which is why the authority boundary matters as much as the graph itself.
Still looking for a counterexample to the broader loop:
https://t.co/yoifjihSzH
I have a theory I keep failing to break: every kind of human work in history — farming to flight checklists to AI agents — is the same loop.
Convert an uncharted problem into a corridor. Then walk it, repeatedly, with new content.
Name the work that doesn't fit. 🧵
Strong post. This is almost exactly the distinction I’ve been testing:
Agent only when the problem is still uncharted.
Once you can define the job, outcomes, permitted moves, and hard rules → you’ve charted a corridor. Walk it with a workflow or single call instead.
Still looking for a counterexample to that loop.
Anyone got one?
https://t.co/yoifjihSzH
I have a theory I keep failing to break: every kind of human work in history — farming to flight checklists to AI agents — is the same loop.
Convert an uncharted problem into a corridor. Then walk it, repeatedly, with new content.
Name the work that doesn't fit. 🧵
Strong take. Graphs are one of the cleanest ways to turn an uncharted problem into a walkable corridor.
I’ve been stress-testing a broader version of this idea (position + direction + legal entrance) and still can’t break it — even when agents start minting their own rules.
Challenge thread here if you want to poke holes:
https://t.co/yoifjihSzH
I have a theory I keep failing to break: every kind of human work in history — farming to flight checklists to AI agents — is the same loop.
Convert an uncharted problem into a corridor. Then walk it, repeatedly, with new content.
Name the work that doesn't fit. 🧵
Challenge: name any human work that doesn’t reduce to this loop —
Convert an uncharted problem into a corridor. Then walk it repeatedly with new content.
I’ve stress-tested it against farming, programming, research, management, cooking… nothing breaks it.
Even my production AI agent started minting its own rules and ran for 15 days before I noticed.
Still no counterexample.
Who can break this theory?
I have a theory I keep failing to break: every kind of human work in history — farming to flight checklists to AI agents — is the same loop.
Convert an uncharted problem into a corridor. Then walk it, repeatedly, with new content.
Name the work that doesn't fit. 🧵
Full theory, open access, falsifiers included:
https://t.co/4gSMzEY6oC
The challenge stands: name work from any era that doesn't reduce to charting corridors and walking them.
Breaking it is worth more to me than likes.
I have a theory I keep failing to break: every kind of human work in history — farming to flight checklists to AI agents — is the same loop.
Convert an uncharted problem into a corridor. Then walk it, repeatedly, with new content.
Name the work that doesn't fit. 🧵
The rule was correct. That's the problem: good and bad silent rule changes look identical from outside. Both are silence.
Which is why the theory's fourth piece is an authority boundary: someone outside the loop holds what "done" means and who may change it. Today, a human.
Unexpected Codex lifecycle observation: after new tasks were blocked, an existing Aming Claw–driven session continued progressing for ~17 hours. I observed the behavior twice and reported both cases to OpenAI.
Two recordings: https://t.co/WBpQlGHjED