Production-ready code can still suffer from performance issues.
Adding custom profiling logs used to be a tedious chore.
And too often, the code became a mess that was harder to reason about.
AI changes the economics of that work.
Code generation is cheap now.
That means I can ask the agent to generate throwaway profiling code:
Add custom application-level events.
Emit structured JSON logs.
Measure the paths I suspect are slow.
Make the output easy to analyze.
Then I can ask the AI to analyze the logs.
This is one of the places where AI is genuinely useful:
summarizing large amounts of structured text.
The agent can figure out the mechanics:
grep, a quick Python script, or whatever else fits the data.
I usually ask for a detailed summary with percentages next to each function or phase.
Then I ask for the bottom line: where should I focus first if I want the highest improvement for the least effort?
Once the first pass gives me a signal, I can add more instrumentation and run the code again.
I have not tried it yet, but this seems like a good use case for agent loops:
ask the agent to add custom profiling logic.
Run the code until it emits enough logs.
Stop and analyze the result.
Decide whether the profiling events should be improved.
When the logs are good enough, the agent can stop the loop. Good enough means the logging is not so noisy that it hurts the measurements.
At that point, the agent should be able to give a rough ranking of the most time-consuming operations.
Then it can give me a final summary with recommendations.
AI can help us find bottlenecks faster by making profiling easier to apply more often.
If this was useful, I would appreciate it if you followed my profile or reposted this.
The riskiest way to use AI on critical code is to ask the agent for one giant change.
The internet is full of posts about agent loop-engineering and loopmaxxing:
how to give agents more autonomy, keep them running longer, and let them continue with less interruption.
That can be useful.
But when I have a delicate coding task, I usually want the opposite.
I want the agent to break the work into smaller moves.
More often than not, the fastest way is to slow down.
That is why I like using what I call the "baby-steps" skill.
The idea is simple:
Make code changes as a sequence of small, reviewable patches, and keep the human in the loop before each one lands.
The workflow is:
1. Pick the smallest coherent next change.
2. Show only that change.
3. Explain the intended effect.
4. Wait for approval.
5. Apply only the approved change.
6. Repeat.
The important part is that approval has to be a real gate.
For every code-editing step, the agent needs an explicit yes.
Silence does not count.
Ambiguity does not count.
Partial agreement does not count.
If I ask for a change to the proposal, the agent revises the proposal and asks again.
Even follow-up fixes need the same discipline.
If an approved change causes formatting or compile issues, the next fix is still a new baby step unless I explicitly broaden the approval.
And the agent is not allowed to hide unrelated edits inside the same step.
So I ask the agent to use a boring, explicit format:
----------------------------------------------
Baby step N
File: target file
Intent: one sentence
Before:
current code
After:
proposed replacement
Approve this baby step before I apply it.
----------------------------------------------
That small ritual is the point.
It turns the agent from "go implement this" into "show me the next safe move."
For critical code, where bugs are expensive, autonomy can be expensive.
Baby steps reduce that risk by slowing the agent down at the right moments.
The agent helps me move faster overall.
I still control the parts that matter: scope, review, and approval.
That is the tradeoff I want for critical code.
Sometimes you want to drive fast.
Sometimes you want to move in baby steps.
The developer drives the agents, and @cosocoio shows them the road.
If this was useful, I would appreciate it if you followed my profile or reposted this.
Code can work and still not be production-ready.
Before I trust a feature, I want to know how it fails.
I like to use AI as someone who is never satisfied and keeps pointing at "what could go wrong?"
I usually ask it the simple questions:
- Do we handle invalid inputs?
- Where can this crash?
- What if the input is valid, but much larger than expected?
- What constraints are we missing?
- Is there a security issue here?
For most production systems, I do not want the program to panic.
I want it to reject bad input, return useful errors, and handle recovery deliberately.
That work matters.
It is also tedious.
Adding validations is repetitive.
Replacing panics with useful errors takes time.
Checking every "impossible" branch takes time.
This is where the agent helps me move faster.
It can suggest guards, better error boundaries, and places where recovery is too optimistic.
But I still have to watch for the usual AI failure mode: redesigning code that only needed a guard.
The developer still owns the decision:
- Should this be an error?
- Should this be an assertion?
- Should this be a validation?
- Should this be a recovery path?
- Or did we just find a bug in our assumptions?
The real value is defensive programming at a lower cost.
AI helps make the boring parts cheaper: validations, error handling, recovery paths, and all the small questions that make code more robust.
If this was useful, I would appreciate it if you followed my profile or reposted this.
High test coverage was always desirable.
The cost was the problem.
AI is changing that equation.
In the previous posts, I wrote about how AI changed the cost of formal verification.
The same thing is happening with tests.
Most codebases have felt this pain:
- A refactor breaks a large number of tests.
- A fixture becomes stale.
- An assertion no longer matches the new behavior.
AI does not make testing free.
But it does reduce the maintenance cost dramatically.
That changes the tradeoff.
If tests are cheaper to write and maintain, we can afford to cover more cases.
This is exactly where I like using agents:
- Finding gaps the current tests do not cover
- Adding another case
- Adjusting fixtures
- Fixing tests after a breaking change
Those tasks are tedious for humans.
For an agent, they are often exactly the right kind of work.
This is the part I try to be careful with.
Cheaper tests are useful.
But they can also create cheaper fake confidence.
I have seen AI write tests that look right at first glance:
they call the function, check that something returns,
and follow the shape of the existing test suite.
But the important question is different:
did the test actually prove the behavior I care about?
I also try to watch for another trap.
AI often treats the current production code as the source of truth.
Most of the time, that is exactly what I want.
But if the production code is wrong, the test can quietly validate the wrong behavior.
The test passes, but my confidence should not.
I see the same pattern when tests break after a change.
AI agents are very good at reading compiler errors and making tests pass again.
That is useful.
But before I accept the fix, I want to know what changed:
did we fix a stale test?
or did we teach the test to accept the wrong behavior?
That is why I try to keep the workflow simple.
First, I want the agent to show me where the current coverage is still weak.
Only then do I want it to write more tests.
1. Ask the agent to find gaps the current tests do not cover.
2. Review the proposed cases before it writes code.
3. Ask it to write the tests.
4. Review the assertions.
5. Ask where invariants or test-only checks would make the code stronger.
6. Make sure the tests check the intended behavior.
AI can make test coverage much cheaper.
That is why I am excited about it.
But I try to keep reminding myself:
The goal is not more tests.
The goal is more confidence.
If this was useful, I would appreciate it if you followed my profile or reposted this.
A verified program is not the finish line.
At the end of the day, the production code is what counts.
In the previous post, I wrote about using AI to turn an English spec into a formally verified program.
I used Dafny, a verification-aware programming language.
Dafny lets you write verifiable programs, and it can also compile them to a few programming languages.
For example, a Dafny spec can be compiled into a Python program.
But the compiled Dafny code is usually not the code I want to run in production.
A Dafny proof does not have to model every production detail.
Sometimes I deliberately remove or simplify parts like parsing, because they distract from the behavior I am trying to verify.
If parsing matters enough, it can get its own focused proof.
Once you move from the verified program to production code, a lot of new concerns enter the picture.
For example:
- Concurrency (threads, futures, mutexes, etc.)
- Persistence (relational database, key-value store, file I/O, etc.)
- Data modeling
- Versioning
- Error handling
- Fault tolerance
- Memory consumption
- Performance
- Observability
- And more
To bridge that gap, I use this workflow:
1. Ask the agent to read the Dafny program.
2. Be explicit about the implementation shape.
For example: APIs, system boundaries, worker threads, and concurrency primitives.
3. Ask the agent to turn the verified program into a concrete implementation plan, starting with happy paths and basic error handling.
4. Ask for mappings in both directions between the Dafny program and the implementation.
If a Dafny invariant says a state must never be reachable, the production code might need an assertion, a helper function, or a comment next to the relevant code.
This step was very helpful for me.
It does not prove the implementation.
But it gives me something concrete to review: where the Dafny rules show up in the production code.
5. Ask the agent to implement one step at a time.
I want to stay in the loop, review the work, and give feedback as the implementation takes shape.
6. Once the implementation is ready, ask for a fresh-eyes review of the Dafny spec and the production code.
If the two are not compatible, ask where the gaps are.
7. Keep iterating: update the implementation, add mappings where required, and return to step 5 or 6 instead of treating the review as a one-time pass.
A formally verified program is not the production implementation.
But it helps find bugs and gaps before the first line of production code is written.
After the implementation starts to take shape, the next challenge is testing it well. AI can help with that too, and I will write about it next.
If this was useful, I would appreciate it if you followed my profile or reposted this.
The best time to catch bugs is before the first line of code.
In the previous post, I wrote about one way AI can improve code quality:
by helping us make the spec less vague before implementation.
This post takes that idea one step further:
using AI not only to clarify the spec, but to help make it verifiable.
Formal verification turns correctness from "I think this works" into something a tool can try to prove.
The spec becomes something a verifier can check, not only something humans can read.
The question changes from "does this seem right?" to "under the rules I wrote, can this system still reach a bad state?"
For a long time, formal verification felt out of reach unless you were already a specialist. AI can help level the playing field.
A few months ago, I read @martinkl's article:
"Prediction: AI will make formal verification go mainstream".
It shifted formal verification in my mind from "interesting idea" to "tool I should actually try."
After reading it, I came across Dafny.
Dafny is a verification-aware programming language.
It lets you write code together with the rules the code must satisfy.
For example: “this must always be true,” or “after this function runs, this result must hold.”
Then it tries to prove the code actually follows those rules.
I am still learning Dafny, and complex specs are not something I can write comfortably on my own yet.
But I can read enough to challenge the result: is this actually saying what I think it is saying?
That alone is a big shift.
When we write an English spec and claim that an invariant should always hold,
we are often still working from intuition.
We often still only have a feeling that it is right.
Dafny turns that feeling into a proof attempt: either Dafny proves the claim, or the Dafny verifier points to the gap in the reasoning.
Once we have a Dafny spec, we can work with the agent to make it verify more meaningful properties:
- Is X always true under these constraints?
- What invariant is missing?
- Does this invariant prove the property we actually care about?
Dafny is one tool in this space. One reason I liked it is that the syntax feels familiar: it looks close to imperative code.
Other tools exist, with different tradeoffs around modeling, verification, and proofs.
In cosoco, I used Dafny to prove correctness for a few parts of the system, especially state invariants in data structures.
For example: can this structure ever reach an invalid state after a sequence of valid operations?
Of course, verifying something in Dafny does not automatically verify the production implementation.
There is still a gap between the verified program and the code that actually runs. That gap is what I want to write about next.
If this was useful, I would appreciate it if you followed my profile or reposted this.
AI slop is real.
The internet is full of examples where AI makes code worse.
But I believe AI can also push us in the opposite direction:
toward clearer thinking, tighter specs, and better code.
The catch is that this only works when we use it with engineering discipline.
For me, one of the strongest examples is spec-first AI coding.
The core idea is simple:
a good spec leaves the agent as little room as possible to guess and make decisions.
That is easier said than done. Perfect precision is hard.
The same flexibility that makes natural language expressive also makes it dangerous as an implementation spec.
The point is not to write the perfect spec immediately.
The point is to keep tightening it until the important decisions are explicit.
This is not always faster than coding manually.
But it changes where the thinking happens.
Instead of discovering the design through generated code, I have to make the important decisions upfront.
Before AI, it was common to start coding and let the design reveal itself along the way.
With AI agents, that gets riskier.
They can move from vague intent to a large implementation before the real design questions have been answered.
The workflow that works best for me:
1. I start by writing the first version of the spec, even if it is rough:
- what the change should do
- what constraints matter
- what invariants must hold
- any implementation direction I already know
2. Then I ask the agent to read the code and grill me on the spec.
The goal is simple: find the ambiguity before it turns into code.
- what decisions are still unspecified
- what assumptions it would need to make
- what edge cases are unclear
- which invariants need stronger wording
3. I answer the questions and update the spec.
4. Then I ask the agent to look for what is still unclear.
5. I repeat this until the spec is precise enough to guide implementation.
Only then do I ask the agent to write code.
But I still do not let it run too far ahead. I ask it to implement in small, reviewable steps.
When the spec gets too large, I do not try to push the whole thing through at once.
I split it into smaller specs, or at least smaller implementation chunks.
That keeps the work easier to review and gives me chances to correct the direction before the agent goes too far.
This is the version of AI coding I trust more:
AI does not improve quality by magically producing perfect code.
It improves quality when it helps me find ambiguity earlier, tighten the spec, and keep the implementation in reviewable pieces.
If the spec is vague, the agent will not magically make it clear.
It will usually turn that vagueness into implementation decisions I did not mean to make.
That is why I want the agent to challenge the spec before it writes code.
If this was useful, I would appreciate it if you followed my profile or reposted this.