LLMs are a gigantic sledgehammer.
Most enterprise failures are not "the model is dumb."
They are someone using a sledgehammer to crack a nut — no sequencing, no tie-break, no thin deterministic glue around the blow.
The model is not the product.
The orchestration that keeps the hammer from destroying the table is.
The naive assumption is that staying in control means reading every line of AI-generated code.
That approach only turns human attention into the bottleneck.
AI changes the unit of control: from the code itself to its observable behavior, constraints, tests, interfaces, and quality gates.
I’m significantly older than you. I started coding in the late 60s. My current strategy is to not read any of the code written by my agents. That’s the only way I can take advantage of their productivity. What I do instead is to surround the agents with extreme constraints. Unit tests, gherkin tests, QA procedures, quality metrics, mutation testing, test coverage, and a plethora of others. In the end, I have very high confidence in the code they produce because they’ve had to run the gauntlet of all of my constraints and tests.
For classical software, the deal was simple: what you can describe, you can automate.
For agents, the deal is uglier: what you can validate, you can automate.
A clever generator without a ruthless critic just produces confident work at higher speed.
The interesting roadmap item is rarely the next tool. It is the judge.
Most office workers do not ask IT to build them a spreadsheet.
They make one.
AI agents will become ordinary when creating one feels closer to opening Excel than requesting a software project.
Evaluation starts with real inputs
We keep trying to evaluate AI systems before we know what real users will ask them.
A synthetic dataset can tell you whether the pipeline works.
It cannot tell you whether the test resembles the job.
TL;DR from 150+ off-stage calls: • job = defend truth under board narrative • ‘show’ beats ‘ship’ in budget theater • headcount is a workshop, not a factory
If you live in feed two, you’re not behind. You’re early to the honest problem.
Which one is loudest where you sit?
After 150+ conversations with Heads of AI, VPs of Eng, and internal champions, I stopped trusting the public AI feed.
Off-stage, the job is almost never ‘ship agents.’
It’s surviving the gap between board story and operational honesty.
5 notes they almost never say on stage:
3/ The people doing the work are not an AI factory.
They’re small squads switching between security questionnaires, data tickets, and Friday’s live demo.
If your strategy assumes a mature platform organization you don’t have, you’re writing fiction.
0.8 × 0.8 × 0.8 × 0.8 × 0.8 = 0.33 - average AI agents accuracy
Your AI agent is 80% accurate at each step.
Put 5 steps into one workflow: 0.8⁵ = 0.33.
Corporate users experience end-to-end reliability—not the best number from your pilot dashboard.