@ApolloResearch A monitor that can actually stop a coding agent is more useful than another screen showing what it did afterwards. I'd test it on one harmless forbidden delete first. If it only explains the danger after the click, it's a dashboard.
OpenAI now rates Astra at 'Critical' for cyber capability. That's an odd product milestone: the model got more useful, and the release conversation immediately became mostly about safeguards. 'More capable' is no longer a simple win. #Astra#Cybersecurity
@danellisona I wouldn't make coding knowledge the entry ticket. Start building, but learn enough to read the important bits and recover when the model gets it wrong. Curiosity can come first. The understanding has to catch up.
@GenkitFramework Stable session stores may be the least exciting line there and the one people feel most. Agents are much easier to trust when they remember the job without you rebuilding the conversation every morning.
@OpenRouter@olam_labs@browser_use@userlens_hq Production demos are where this gets useful. Ask every founder the same boring question: what broke after the happy-path demo, and which model change actually fixed it?
@dgalarza Typed tools beat making an agent squint at buttons. I'm curious how much of Shopify is exposed here. Can an agent actually manage a cart and checkout, or is this mostly product discovery for now?
@karch_andreas@AnthropicAI Giving an agent three free months and letting it install numerical tools is a magnificent way to learn. I suspect this thread is about to explain why 'run amok' was doing quite a lot of work.
@notrab Eight billion database operations because `<` should've been `<=` is brutal. Also a perfect reminder that the most expensive bugs can look hilariously small in the diff. Fair play for publishing the whole post-mortem.
@Jadu100x The number one lie is usually followed by 'while I'm here, I may as well add auth'. Three hours later you're debugging email verification and the original button is still wonky.