@GoogleDeepMind “Autonomous patching” is where the word autonomous needs a footnote. Finding the fix can be automated; shipping it still needs scope, a stop button and a rollback path.
@cursor_ai Self-hosted changes the perimeter, not the trust problem. If the agent can reach internal services and custom hardware, every call still needs a boundary.
@SpaceXAI A good eval shouldn’t stop at whether the prompt was refused. The real blast radius starts after that: tools, files, credentials and outbound actions.
@openclaw Bug fixes are painful. Permission changes are worse. When an agent is wired into real tools, every release should answer one blunt question: what can it reach now?
@albertwenger Call them goals, objectives or learned behavior, the security question doesn’t change. If the system has tools and authority, what can it touch, when must it stop, and who can pull the plug? Semantics won’t save a bad permission model.
@Da7_Tech Everyone’s arguing model vs. UX, but the failure mode is simpler: the agent keeps the keys while it loses the plot. Give it a narrow task, a hard stop before the next tool call, and a clean recovery path. Otherwise every “fix” just buys you another loop.
@AlexFinn OpenClaw 2.0 looks like a product now, which makes the reliability gap harder to excuse. A slick UI won’t save an agent that loops, wanders or reaches for the wrong tool. The boring bits scope, stop conditions and recovery are still the product.
@VaibhavSisinty 13,000 skills is impressive. It also changes the question from “can the agent do this?” to “what exactly did I just install?” I’d want to see the tools, the data they can reach and an easy way to revoke them. Otherwise it’s npm with a cape.
@mitsuhiko Being the subagent is funny until it can write the firmware instead of just asking you to replug the USB. Hardware access needs the same discipline as cloud tools: clear capabilities, limited actions and a real stop button when the agent starts wandering
@openclaw The install experience is the easy win. The harder question starts after the skill is installed: what did it add, what can it touch, and can I take it away without digging through the whole system? An agent OS needs app permissions, not just an app store.
@dhh The moment people get agent-pilled, the permission question arrives right behind it. Read, write, send, spend different powers, different locks. “It keeps working” is the feature; knowing exactly where it must stop is the product.
@celineodier Giving an agent eyes is the easy part. What happens when those eyes can also write, send, publish or spend? Every new repo adds a capability—and a new place where the agent should have to stop and ask.
@monokern Five bots sharing one browser, terminal and authenticated tools aren’t five identities. It’s one shared blast radius with five interfaces. The real feature is the permission boundary: a specialist can help without inheriting every credential its neighbors can reach.
@mattpocockuk “The agent did it, not me” is funny right up until the action hits prod. Once it crosses terminal, browser, and integrations, responsibility can’t vanish at the handoff the call still needs scope, a stop condition, and a trace.
@dylanpkel Pay-per-use capabilities are a compelling primitive. The hard question is who can authorize a new capability mid-task and what stops an agent’s budget from quietly becoming permission to do anything?
@calcsam On-the-fly skill loading is a UX win and a new trust boundary. Before a skill becomes callable, I’d want to see the tools it introduces, the data it can touch, and when those permissions expire.
@mattpocockuk Shrinking the system prompt changes what the agent sees; disabling tools changes what it can reach. They solve different risks. I’d still check the concrete call before execution shorter prompts don’t make dangerous actions safe.