This is the most annoying thing that I cannot solve. I have all sort of rules in different places but agents will eventually commit and push stuff I don't tell them to. Maybe something in the system prompt?
On the other hand somewhat of a reality check. Even frontier AI can't a follow a simple instruction.
Only in Germany: police declares manhunt on "armed and dangerous" suspect. Asks citizens to call if they see him, but can't share picture because of GDPR.
@nida_banou@BrennpunktUA Ein Foto finden Sie unter folgendem Link: https://t.co/gzhjx9qVx6
Aus datenschutzrechtlichen Grรผnden kรถnnen wir dies nicht direkt auf Social Media teilen.
NEW: San Francisco woman claims her husband handed her nearly all parenting duties so he could become โAI nativeโ โ spending days, nights, & weekends locked in his office learning AI.
> In other words, we find the last 5% (ie where a model is truly equivalent to a reasonable engineer) extremely difficult to achieve
When will we learn?
8090 works on production systems for large, often regulated, enterprises.
Vibing isnโt tolerated because these are the systems that run western society - banking, power, healthcare, insurance etc.
Over the last few quarters, the gains that we got from using frontier models inside of our Software Factory on these systems started to shrink but the costs kept doubling. This makes sense I guess, as in hindsight, we were initially asking the model to do mostly light work (generate basic PRs) and now we were asking it to do more complex work (mitigate dependencies across systems).
Unless you grow context massively, be willing to run many A/B tests and iterate massively (ie use massively more tokens) complex tasks stay roughly unfinished by the model and requires the engineer to largely act alone.
In other words, we find the last 5% (ie where a model is truly equivalent to a reasonable engineer) extremely difficult to achieve and extremely expensive to such a degree that the fully loaded cost of the model + the engineer will not pay for itself.
So I asked our CTO to start thinking about other ways. We need our engineers to have access to the best tools BUT we also need to educate them to think even more for themselves - not less - in this last mile.
At the same time, we need to find solutions that decrease our token costs by 90% - especially because these bleeding edge tokens are not nearly as cost effective as the tokens before it and are creating a big OpEx bill for us.
I wonder how many engineers, in all orgs, are running amok right now by using the latest frontier models as a kind of slot machine. Increasingly turning their mind off, largely keeping productivity flat while their CEO and CFO deals with a massive token bill?
My advice to you is that when you encounter this last 5% of very hard technical challenges in getting a complex system into production, be circumspect.
The challenge of the last 5% is actually getting harder - especially as hundreds and thousands of code generation model runs run amok adding all kinds of random cruft into codebases that eventually need to be rationalized.