It’s insane to me how much the tiniest bit of steering/context engineering improves agentic engineering outputs.
E.g. when building an agentic feature, simply telling Opus “give the agent minimal instructions, give it a virtual filesystem, and let it be autonomous” can produce dramatically better results than not saying it.
That bit of understanding is the difference maker between great results and avg results.
@GergelyOrosz +1. I've been through several rounds of it. It ends up task A has a semi-blocker, task B also has a semi-blocker, task C has a semi-blocker, no wonder these things are in my backlog.
I swear the most annoying thing in the world is being forced to call Fidelity to manage your finances just so they can try to sell you advisory services. Please I beg you, let us manage via web/app.
As usual, thanks for the writeup. Especially after reading your thoughts I have to agree that the long-horizon slider was maxed, too quickly. It was RL'd absurdly, and to your point, found quite strange ways to get its carrots. They're still too fallible for it to be practical to unleash for several hours without oversight, at least for SWE things that benefit from such oversight. I understand your point that maybe they're for someone else but I don't really think that customer exists as much as implied.
@thsottiaux@oneill_c Well yeah, but surely you can see the dissonance between this reality and ‘coding is solved’ type of rhetoric. Anyways, cool to see that even these more bleeding edge / niche things are benefiting more from agents.
@GergelyOrosz *And because it’s their own product, lol. I’m sure cursor staff is cursor pilled, anthropic is CC pilled, etc.. this is not a meaningful insights.