IBM CTO dropped a guide on letting AI run:
00:00 - the black box: an AI decides, no explanation, no appeal
01:15 - plan-do-check-act: running AI as a continuous loop, not a one-off
02:00 - the 4 layers where it actually goes wrong: data, model, system, use
03:10 - who owns it: if no one's responsible, it doesn't get done
letting an agent run for you is powerful, but only if someone still owns what it does.
how to build one you can trust is in the article below.
Microsoft just read ~300 papers on self-improving agents and found the one thing that separates a loop that gets better from a loop that quietly rots.
propose a change → let something you can't fool check it → keep what passes → repeat
when the thing grading the work is the system doing the work, the score drifts - the loop optimizes the number, not the job, and the longer it runs the worse it gets.
autonomous loops only kept improving where the verifier was deterministic and independent of the agent.
the self-graded ones degraded every iteration.
the writer proposes. something the writer can't be decides if it's done.
read the paper first, then the article below.
Microsoft just read ~300 papers on self-improving agents and found the one thing that separates a loop that gets better from a loop that quietly rots.
propose a change → let something you can't fool check it → keep what passes → repeat
when the thing grading the work is the system doing the work, the score drifts - the loop optimizes the number, not the job, and the longer it runs the worse it gets.
autonomous loops only kept improving where the verifier was deterministic and independent of the agent.
the self-graded ones degraded every iteration.
the writer proposes. something the writer can't be decides if it's done.
read the paper first, then the article below.
OpenAI:
"every time you have to interact with the agent is a failure of the harness."
if you're babysitting it, the setup is doing its job wrong.
a good loop runs, checks itself, and pings you only when it's stuck.
that setup is in the article below.
@oaktoebark hard to argue when the trigger is literally "institution launches chain, memecoin pumps"
doesnt even need a narrative anymore, just a new chain name