I’ve been working on the same problem. Doing a static pressure test of various scenarios and their outcomes to find ones that are critical in moving closer to a given goal.
The best analogy to this is chess. All of the existing pieces have their roles and constraints. You can execute any combination of moves to reach a specific outcome. Whether that is to win the match (long goal) or it is to create the conditions to improve my chances to win (short goal) depends on a lot of variables.
Strategy for an agent can be as simple as just identifying the broader conditions that prevent some combination of tasks or tickets from completing. Choosing what to do with that is harder still.
It's basically an attempt to turn an agent into a strategic thinker. So far it feels really fucking hard and will likely turn the agent into a slop cannon in most cases.
But always worth trying something that challenges your priors.
@mattpocockuk@skastr052 It is. But not all code is worth making testable. Sometimes the answer is to just start again with lessons learned. The human skill to develop is knowing when and why to make that choice.
@CognitiveTake@mattpocockuk He’s saying people shouldn’t care ONLY about better models. The quality comes from what you invest into the tools that manipulate the models.
Maybe you missed the subtext?
Claude Mods are landing now. Someone already built a Tetris-in-Claude mod 🤯
See issue for the latest community update, technical details, and more cool demos
https://t.co/A15qGUZ6nx
I’ve been using Opus 5 for a while now. It’s not great. Its output style makes simple ideas and statements deeply confusing and seems optimized for unmanaged use. I switched back to Opus 4.6 and what a difference it makes.
I don’t believe it’ll be a problem for Cursor. The model helps but it isn’t what made Cursor Compose great. They’ll do just fine with any competent model whose performance resembles anything from a year or more ago. Open models will catch up and this will not have mattered.