@FUCORY I'd have the harness record which task created each tab or worktree, then check for unfinished changes and other tasks still using it. That would give the user a concrete cleanup list to approve.
@JoshARosen Does the skill need to contain the decision logic? For something like lead prioritization, it could call a decision API with the current account context, while the proprietary criteria stay server-side.
@dair_ai Would be useful to compare full tool schemas vs on-demand loading in the same harness. Does the smaller context still save money once you count discovery calls and extra steps?
@eyad_khrais Approval should be enforced by the tool. A skill can change how the agent handles a case, but editing its instructions shouldn’t let it bypass approval.
@dabit3 I’d test a skill against specific tasks before keeping or deleting it. Reading a PDF, quoting a passage, and reusing a figure need different instructions. A skill can help with one and add noise to another.
@0xDevShah I'd freeze the model and test context/routing changes on held-out workflows first. If the same reasoning failures survive that, there's a much clearer case for owning the model.
@joseemv88 How do you handle rules like “material scope changes need client sign-off”? The units conflict can be compiled. I'd keep the scope-change judgment as an explicit, tested decision, with the approval requirement enforced separately at execution.
@HamelHusain For the handoff case, I'd label the decision point: given the context available then, should the agent act, ask, or hand off? Scoring only the final answer can hide the wrong branch. Each failure class needs its own cases before the agent starts optimizing it.