@BHolmesDev@poteto the outer loop grades how the agent worked. does anything grade what it shipped? "monitor usage" is in your inner loop. does it ever feed a skill change, or only new issues?
@not_ai@BHolmesDev usage logs, yeah. unused -> prune is one case. the other is drift: the line still passes but users moved on from it. do your e2e titles ever get rewritten from usage, or only when someone files a change?
@housecor the CLI point is the one I'd push on. if the agent runs the check and reports the result, it's still grading itself. does the factory run that CLI independently, or trust what the agent says it saw?
@not_ai@BHolmesDev the constitution diff is the part I'd steal. do you ever run it the other direction, find lines that no longer match what users actually do, and delete them?
the studio's job is judgment: what to make, why, and what to kill. a publisher turns down most of what it gets. a faster press doesn't pick better books, and more context doesn't make a factory a studio.
software factories are missing an abstraction. people keep trying to turn the factory into the product studio. that's a printing press trying to become a publisher.
trust in software is gate-shaped: prove at design, review at merge, test at commit. gates check code against the world in the moment. then traffic 10x's, users extend features, a dep changes. "correct" changes. verification is a control loop, not a gate.
@zachlloydtweets the missing piece between this and "terraform for factories" is the reconciliation loop. terraform provisions, k8s converges (desired state, observed state). agents are pods, not the platform.
and so every factory needs two loops. inner-loop reconciles the system to its declared invariants, outer-loop reconciles invariants to the world and its outcomes.
trust in software is gate-shaped: prove at design, review at merge, test at commit. gates check code against the world in the moment. then traffic 10x's, users extend features, a dep changes. "correct" changes. verification is a control loop, not a gate.
a factory needs two measures. one watches the line: are the steps between intake and shipped code drifting from what we qualified? one watches the world: are the outcomes still right? all our tooling is the first kind. the chart reads green while the product is wrong.
1/ When we looked at the economics of AGI, the key policy challenge was immediately clear:
AI drastically lowers the cost of execution for anything easy to verify.
For everything else, verification is the bottleneck.
@JoshARosen@typesafeai 💯 the drift + verification split is a useful decomposition. curious about the judge side: when foreman assesses "verification," is that deterministic checks it runs, or a model grading the work?
skills get read, not followed: agents retrieve the right instruction ~96% of the time, follow it ~35% (arxiv 2605.30621). the pattern i kept seeing and validated locally: skills that structure the workflow get followed. skills that advise get read.