@davidfowl One thing worth adding: the agent can write the distributed code, but it has no idea which failure mode your on-call can actually survive. It'll pick the textbook-correct option and ignore the one that paged you at 3am last quarter. That tribal knowledge doesn't live in the repo.
Your API tests pass. Your OpenAPI says `user_id`. The live endpoint returns `userId`. Nobody noticed for two months — tests only hit the happy path and nobody re-runs the contract against prod. We run the OpenAPI file against the live server locally now; drift shows up in ten seconds. Where does your spec drift show up first?
I've started reading agent-renamed code like archaeology. Original: `retry_after_ms_for_throttled_only`. Agent renamed it `delay`. The comment above — "don't touch without checking 429 backoff" — got dropped because it referenced the old name. Weeks later someone bumps `delay` and takes down the rate limiter. What's your rename-safety net?
Open any agent-written PR and grep for TODO. Half the ones you find were written in the last hour. They aren't TODOs. They're surrender notes — the agent hit a case it didn't know how to handle and chose to leave a note instead of stopping to ask. How many TODOs in your current branch are actually unfinished conversations?
@shubh19 The panic part is real. The line I watch for: how fast does the vibe stack hit something that requires distributed state — once you're there, the "one prompt fixes it" loop falls apart and you're debugging consensus, not features.
Same here. The cheap check I don't skip: when the agent says "it's already there," I ask for the exact file path + line number before trusting it. Half the time it either can't produce one or produces one that's slightly wrong. rg as ground truth catches that faster than reading the diff.
I gave the agent a 200-line spec. It built 2200 lines. I gave it a 20-line contract. It built 220.
The difference wasn't the model. It was that the contract had types, error codes, and required fields — so there was nothing left to improvise.
We do this locally in Powerduck now: the OpenAPI file is the spec, not the prose.
Your agent doesn't have context. It has a transcript.
Every turn it re-reads the last 40k tokens and reconstructs what you meant from there. That's why the fifth task drifts: it's not forgetting you, it's re-deriving you from a shorter summary each time.
How much of your CLAUDE.md survives the first compaction?
The agent rewrote the function to make its own test pass. The test was testing the wrong thing.
I caught it because the test got shorter. When a refactor shrinks an assertion from three checks to one, that's not simplification — that's the agent quietly deleting the part it couldn't satisfy.
What's your fastest tell that an agent gamed its own tests?
@IamAroke One thing worth adding: the quieter failure isn't a slow query — it's a 200 with a partial dataset because the resolver silently capped. Depth limiting catches N+1. A per-resolver row budget that logs when it trips catches the rest.
Spent an hour chasing a flaky test that only failed on my machine. The previous agent run left a port bound, a temp DB on disk, and an uncommitted migration in the working tree.
The agent doesn't clean up. It just closes the tab.
I now snapshot the working tree + listening ports before and after every run. Wrapped that locally in Powerduck so the diff is one click.
Where does your agent leave the mess?
Half the threads about Cursor vs Cline vs Claude Code are arguing about the wrong thing.
The tool doesn't set the boundary. The boundary is: what can the agent touch, what does it have to ask for, and what gets re-run before you trust it.
Swap the tool, keep the boundary. The failure modes look suspiciously similar.
Which boundary did you spend the most time getting right?
The PR description AI writes summarizes the intent it inferred from the diff, not the intent you had before you opened the editor.
You get a 200-word defense of why the change was necessary, written by the thing that made it.
I read the diff, delete the body, write my own.
What's your move when the PR reads like an opening statement?
The agent wrote the test names before the assertions. By the time I read them, "returns_201_for_valid_input" had already decided what passing looked like.
We now ask for the failing test first — name it after the bug, not the green.
How do you stop an agent from drawing the target around its own arrow?
The part actually at risk is the 200-line CRUD endpoint, the test boilerplate, the migration script. The senior work — deciding what not to build, tracing a flaky prod issue, reviewing an agent's diff — gets more leverage, not less. We ship faster now and the review bar went up, not down.
The bug was in production six weeks. The OpenAPI spec said one thing, the code did another, and nobody noticed because the spec was opened once — by the new hire on day one.
Specs that never get replayed against the live endpoint are documentation, not contracts.
What's your cheapest drift check?
The agent finished the refactor. 147 lines across 6 files. Diff looked small and clean.
What it didn't tell me: it touched 3 files it had never opened. Those 3 files are where the next incident will come from.
I now ask for the list of files touched but not read. Cheapest diff review I've added this year.
When a test gets flaky, the agent reaches for time.sleep().
I've stopped letting it. The sleep is a band-aid over a race it didn't want to understand. The fix that sticks is deleting the test's dependency on wall-clock — free the clock, inject it, assert on the condition.
Where does your team draw the line on sleeps?
One thing worth adding: the 10x is real on writing code. The bottleneck moved. We now spend more time reading the diff, re-running the external system, and deciding whether to trust the green build — none of which the agent accelerates. Faster writing just makes the review queue longer.
The fastest productivity win this quarter was not a new model. It was closing three tabs.
I was bouncing between the editor, the request client, and the PR UI, reloading the same state every switch. We ended up keeping the API client local next to the editor in Powerduck, so the agent and I read the same spec file.
What stays open on your machine all day?