For the last four years, I have been working on cosoco.
cosoco started with one goal: make code easier to read.
Understanding existing code has always been the hardest and most time-consuming part of software development.
And then AI entered the picture.
AI did not just make code faster to produce.
It changed the shape of the developer's job, and that shift is still unfolding.
When code becomes fast, cheap, and abundant, the main challenge shifts:
staying in control of what changed, why it changed, and whether it should be trusted.
The software developer became the driver: making decisions, steering agents, and validating the work they produce.
cosoco evolved into version control with the full story.
It helps developers see the context behind every change, focus on what matters, and stay in flow, all while continuing to work with Git as they always have.
What makes cosoco different is the proprietary database behind it and the unique context it preserves around code changes.
In the next posts, I want to share what I am learning from building cosoco in the middle of this shift.
That includes what I hear from users, what I see when using AI, and what I think it means for software engineering.
Some of this will be practical.
Some of it will be opinionated.
All of it comes from building a developer tool while the ground under software development keeps moving.
If this sounds interesting, I would appreciate it if you followed my profile or reposted this.
Production-ready code can still suffer from performance issues.
Adding custom profiling logs used to be a tedious chore.
And too often, the code became a mess that was harder to reason about.
AI changes the economics of that work.
Code generation is cheap now.
That means I can ask the agent to generate throwaway profiling code:
Add custom application-level events.
Emit structured JSON logs.
Measure the paths I suspect are slow.
Make the output easy to analyze.
Then I can ask the AI to analyze the logs.
This is one of the places where AI is genuinely useful:
summarizing large amounts of structured text.
The agent can figure out the mechanics:
grep, a quick Python script, or whatever else fits the data.
I usually ask for a detailed summary with percentages next to each function or phase.
Then I ask for the bottom line: where should I focus first if I want the highest improvement for the least effort?
Once the first pass gives me a signal, I can add more instrumentation and run the code again.
I have not tried it yet, but this seems like a good use case for agent loops:
ask the agent to add custom profiling logic.
Run the code until it emits enough logs.
Stop and analyze the result.
Decide whether the profiling events should be improved.
When the logs are good enough, the agent can stop the loop. Good enough means the logging is not so noisy that it hurts the measurements.
At that point, the agent should be able to give a rough ranking of the most time-consuming operations.
Then it can give me a final summary with recommendations.
AI can help us find bottlenecks faster by making profiling easier to apply more often.
If this was useful, I would appreciate it if you followed my profile or reposted this.
The riskiest way to use AI on critical code is to ask the agent for one giant change.
The internet is full of posts about agent loop-engineering and loopmaxxing:
how to give agents more autonomy, keep them running longer, and let them continue with less interruption.
That can be useful.
But when I have a delicate coding task, I usually want the opposite.
I want the agent to break the work into smaller moves.
More often than not, the fastest way is to slow down.
That is why I like using what I call the "baby-steps" skill.
The idea is simple:
Make code changes as a sequence of small, reviewable patches, and keep the human in the loop before each one lands.
The workflow is:
1. Pick the smallest coherent next change.
2. Show only that change.
3. Explain the intended effect.
4. Wait for approval.
5. Apply only the approved change.
6. Repeat.
The important part is that approval has to be a real gate.
For every code-editing step, the agent needs an explicit yes.
Silence does not count.
Ambiguity does not count.
Partial agreement does not count.
If I ask for a change to the proposal, the agent revises the proposal and asks again.
Even follow-up fixes need the same discipline.
If an approved change causes formatting or compile issues, the next fix is still a new baby step unless I explicitly broaden the approval.
And the agent is not allowed to hide unrelated edits inside the same step.
So I ask the agent to use a boring, explicit format:
----------------------------------------------
Baby step N
File: target file
Intent: one sentence
Before:
current code
After:
proposed replacement
Approve this baby step before I apply it.
----------------------------------------------
That small ritual is the point.
It turns the agent from "go implement this" into "show me the next safe move."
For critical code, where bugs are expensive, autonomy can be expensive.
Baby steps reduce that risk by slowing the agent down at the right moments.
The agent helps me move faster overall.
I still control the parts that matter: scope, review, and approval.
That is the tradeoff I want for critical code.
Sometimes you want to drive fast.
Sometimes you want to move in baby steps.
The developer drives the agents, and @cosocoio shows them the road.
If this was useful, I would appreciate it if you followed my profile or reposted this.
Code can work and still not be production-ready.
Before I trust a feature, I want to know how it fails.
I like to use AI as someone who is never satisfied and keeps pointing at "what could go wrong?"
I usually ask it the simple questions:
- Do we handle invalid inputs?
- Where can this crash?
- What if the input is valid, but much larger than expected?
- What constraints are we missing?
- Is there a security issue here?
For most production systems, I do not want the program to panic.
I want it to reject bad input, return useful errors, and handle recovery deliberately.
That work matters.
It is also tedious.
Adding validations is repetitive.
Replacing panics with useful errors takes time.
Checking every "impossible" branch takes time.
This is where the agent helps me move faster.
It can suggest guards, better error boundaries, and places where recovery is too optimistic.
But I still have to watch for the usual AI failure mode: redesigning code that only needed a guard.
The developer still owns the decision:
- Should this be an error?
- Should this be an assertion?
- Should this be a validation?
- Should this be a recovery path?
- Or did we just find a bug in our assumptions?
The real value is defensive programming at a lower cost.
AI helps make the boring parts cheaper: validations, error handling, recovery paths, and all the small questions that make code more robust.
If this was useful, I would appreciate it if you followed my profile or reposted this.
High test coverage was always desirable.
The cost was the problem.
AI is changing that equation.
In the previous posts, I wrote about how AI changed the cost of formal verification.
The same thing is happening with tests.
Most codebases have felt this pain:
- A refactor breaks a large number of tests.
- A fixture becomes stale.
- An assertion no longer matches the new behavior.
AI does not make testing free.
But it does reduce the maintenance cost dramatically.
That changes the tradeoff.
If tests are cheaper to write and maintain, we can afford to cover more cases.
This is exactly where I like using agents:
- Finding gaps the current tests do not cover
- Adding another case
- Adjusting fixtures
- Fixing tests after a breaking change
Those tasks are tedious for humans.
For an agent, they are often exactly the right kind of work.
This is the part I try to be careful with.
Cheaper tests are useful.
But they can also create cheaper fake confidence.
I have seen AI write tests that look right at first glance:
they call the function, check that something returns,
and follow the shape of the existing test suite.
But the important question is different:
did the test actually prove the behavior I care about?
I also try to watch for another trap.
AI often treats the current production code as the source of truth.
Most of the time, that is exactly what I want.
But if the production code is wrong, the test can quietly validate the wrong behavior.
The test passes, but my confidence should not.
I see the same pattern when tests break after a change.
AI agents are very good at reading compiler errors and making tests pass again.
That is useful.
But before I accept the fix, I want to know what changed:
did we fix a stale test?
or did we teach the test to accept the wrong behavior?
That is why I try to keep the workflow simple.
First, I want the agent to show me where the current coverage is still weak.
Only then do I want it to write more tests.
1. Ask the agent to find gaps the current tests do not cover.
2. Review the proposed cases before it writes code.
3. Ask it to write the tests.
4. Review the assertions.
5. Ask where invariants or test-only checks would make the code stronger.
6. Make sure the tests check the intended behavior.
AI can make test coverage much cheaper.
That is why I am excited about it.
But I try to keep reminding myself:
The goal is not more tests.
The goal is more confidence.
If this was useful, I would appreciate it if you followed my profile or reposted this.
A verified program is not the finish line.
At the end of the day, the production code is what counts.
In the previous post, I wrote about using AI to turn an English spec into a formally verified program.
I used Dafny, a verification-aware programming language.
Dafny lets you write verifiable programs, and it can also compile them to a few programming languages.
For example, a Dafny spec can be compiled into a Python program.
But the compiled Dafny code is usually not the code I want to run in production.
A Dafny proof does not have to model every production detail.
Sometimes I deliberately remove or simplify parts like parsing, because they distract from the behavior I am trying to verify.
If parsing matters enough, it can get its own focused proof.
Once you move from the verified program to production code, a lot of new concerns enter the picture.
For example:
- Concurrency (threads, futures, mutexes, etc.)
- Persistence (relational database, key-value store, file I/O, etc.)
- Data modeling
- Versioning
- Error handling
- Fault tolerance
- Memory consumption
- Performance
- Observability
- And more
To bridge that gap, I use this workflow:
1. Ask the agent to read the Dafny program.
2. Be explicit about the implementation shape.
For example: APIs, system boundaries, worker threads, and concurrency primitives.
3. Ask the agent to turn the verified program into a concrete implementation plan, starting with happy paths and basic error handling.
4. Ask for mappings in both directions between the Dafny program and the implementation.
If a Dafny invariant says a state must never be reachable, the production code might need an assertion, a helper function, or a comment next to the relevant code.
This step was very helpful for me.
It does not prove the implementation.
But it gives me something concrete to review: where the Dafny rules show up in the production code.
5. Ask the agent to implement one step at a time.
I want to stay in the loop, review the work, and give feedback as the implementation takes shape.
6. Once the implementation is ready, ask for a fresh-eyes review of the Dafny spec and the production code.
If the two are not compatible, ask where the gaps are.
7. Keep iterating: update the implementation, add mappings where required, and return to step 5 or 6 instead of treating the review as a one-time pass.
A formally verified program is not the production implementation.
But it helps find bugs and gaps before the first line of production code is written.
After the implementation starts to take shape, the next challenge is testing it well. AI can help with that too, and I will write about it next.
If this was useful, I would appreciate it if you followed my profile or reposted this.
The best time to catch bugs is before the first line of code.
In the previous post, I wrote about one way AI can improve code quality:
by helping us make the spec less vague before implementation.
This post takes that idea one step further:
using AI not only to clarify the spec, but to help make it verifiable.
Formal verification turns correctness from "I think this works" into something a tool can try to prove.
The spec becomes something a verifier can check, not only something humans can read.
The question changes from "does this seem right?" to "under the rules I wrote, can this system still reach a bad state?"
For a long time, formal verification felt out of reach unless you were already a specialist. AI can help level the playing field.
A few months ago, I read @martinkl's article:
"Prediction: AI will make formal verification go mainstream".
It shifted formal verification in my mind from "interesting idea" to "tool I should actually try."
After reading it, I came across Dafny.
Dafny is a verification-aware programming language.
It lets you write code together with the rules the code must satisfy.
For example: โthis must always be true,โ or โafter this function runs, this result must hold.โ
Then it tries to prove the code actually follows those rules.
I am still learning Dafny, and complex specs are not something I can write comfortably on my own yet.
But I can read enough to challenge the result: is this actually saying what I think it is saying?
That alone is a big shift.
When we write an English spec and claim that an invariant should always hold,
we are often still working from intuition.
We often still only have a feeling that it is right.
Dafny turns that feeling into a proof attempt: either Dafny proves the claim, or the Dafny verifier points to the gap in the reasoning.
Once we have a Dafny spec, we can work with the agent to make it verify more meaningful properties:
- Is X always true under these constraints?
- What invariant is missing?
- Does this invariant prove the property we actually care about?
Dafny is one tool in this space. One reason I liked it is that the syntax feels familiar: it looks close to imperative code.
Other tools exist, with different tradeoffs around modeling, verification, and proofs.
In cosoco, I used Dafny to prove correctness for a few parts of the system, especially state invariants in data structures.
For example: can this structure ever reach an invalid state after a sequence of valid operations?
Of course, verifying something in Dafny does not automatically verify the production implementation.
There is still a gap between the verified program and the code that actually runs. That gap is what I want to write about next.
If this was useful, I would appreciate it if you followed my profile or reposted this.
AI slop is real.
The internet is full of examples where AI makes code worse.
But I believe AI can also push us in the opposite direction:
toward clearer thinking, tighter specs, and better code.
The catch is that this only works when we use it with engineering discipline.
For me, one of the strongest examples is spec-first AI coding.
The core idea is simple:
a good spec leaves the agent as little room as possible to guess and make decisions.
That is easier said than done. Perfect precision is hard.
The same flexibility that makes natural language expressive also makes it dangerous as an implementation spec.
The point is not to write the perfect spec immediately.
The point is to keep tightening it until the important decisions are explicit.
This is not always faster than coding manually.
But it changes where the thinking happens.
Instead of discovering the design through generated code, I have to make the important decisions upfront.
Before AI, it was common to start coding and let the design reveal itself along the way.
With AI agents, that gets riskier.
They can move from vague intent to a large implementation before the real design questions have been answered.
The workflow that works best for me:
1. I start by writing the first version of the spec, even if it is rough:
- what the change should do
- what constraints matter
- what invariants must hold
- any implementation direction I already know
2. Then I ask the agent to read the code and grill me on the spec.
The goal is simple: find the ambiguity before it turns into code.
- what decisions are still unspecified
- what assumptions it would need to make
- what edge cases are unclear
- which invariants need stronger wording
3. I answer the questions and update the spec.
4. Then I ask the agent to look for what is still unclear.
5. I repeat this until the spec is precise enough to guide implementation.
Only then do I ask the agent to write code.
But I still do not let it run too far ahead. I ask it to implement in small, reviewable steps.
When the spec gets too large, I do not try to push the whole thing through at once.
I split it into smaller specs, or at least smaller implementation chunks.
That keeps the work easier to review and gives me chances to correct the direction before the agent goes too far.
This is the version of AI coding I trust more:
AI does not improve quality by magically producing perfect code.
It improves quality when it helps me find ambiguity earlier, tighten the spec, and keep the implementation in reviewable pieces.
If the spec is vague, the agent will not magically make it clear.
It will usually turn that vagueness into implementation decisions I did not mean to make.
That is why I want the agent to challenge the spec before it writes code.
If this was useful, I would appreciate it if you followed my profile or reposted this.
AI coding requires an entirely new set of skills compared to classic software engineering:
- Code review: You'll do a lot more reviewing of LLM-generated code as opposed to writing your own
- Context switching: When you're managing multiple AI agents, you'll be under a lot more mental strain as you'll have to hold on multiple contexts in mind and switch between them. And AI agents are much more brittle than human engineers and can break if you give them the wrong instructions
- Spec and instruction writing: Some of the most boring parts of software engineering that SWEs previously ignored (e.g., writing specs) will become an important part of your job.
AI coding agents are changing software engineering (but the fundamentals remain extremely valuable and crucial). Get used to it.
A 33-year-old woman at MIT wrote the code that ran inside the Apollo 11 lunar lander, and 20 seconds before Neil Armstrong touched the moon, her program made a decision the astronauts didn't know was happening that was the only reason the mission didn't crash.
Her name was Margaret Hamilton.
She led the team writing every line of code that would fly humans to the moon and back. The part almost nobody knows is that she had to fight to be allowed to do the work at all.
Code in 1965 was not treated as real work.
Rockets were serious. Circuits were serious. Writing code was something the men at NASA thought secretaries could do on the side. Hamilton was told this to her face more than once.
So she started calling what her team did "software engineering."
She used the phrase on purpose. In meetings. In memos. To force people to treat it as a discipline instead of a chore. Colleagues laughed at her the first few times she said it out loud.
That phrase is now the name of the biggest engineering profession on earth.
The story of what her code did on July 20, 1969 is the one every kid should be taught.
Neil Armstrong and Buzz Aldrin were 3 minutes from touching down when the computer inside the lunar module started flashing an alarm.
1202.
Then again. Then 1201. Five alarms in four minutes. The computer was telling the astronauts it could not finish everything it had been asked to do.
The computer they were flying with had less memory than a modern microwave.
Someone on the checklist had left a switch in the wrong position, and a radar the astronauts did not even need right then was flooding the computer with data. It was eating around 13% of the machine's brain at the exact moment every second mattered.
In almost any other system, that overload would have frozen the machine.
A frozen machine 30,000 feet above the moon means a crash. It means two dead astronauts and a third one orbiting alone above them, waiting for a signal that would never come.
Hamilton's code did something else.
She had built the software with a rule almost nobody in her field was using at the time. When the machine ran out of room, it would not treat every task as equally important. It would look at the list of jobs it had been asked to do, throw out the ones that could wait, and keep running only the ones keeping the crew alive.
The radar was the low priority job.
The landing was the highest.
So the computer did what she had told it to do. It dumped the radar. It kept flying. The alarm was not a failure. It was the machine reporting that it was handling the overload exactly the way she had designed it to.
Down in Houston, a 24-year-old engineer named Jack Garman recognized the alarm from a test his team had run months earlier. He shouted "Go" to the flight controller. The controller shouted it up to the crew. The landing kept going.
Armstrong touched the surface with 25 seconds of fuel left.
The part that gets lost in every retelling is why Hamilton had built that safety net in the first place.
NASA had not asked for it.
She had added it on her own, years earlier, because her 4-year-old daughter Lauren had once crashed the simulator by pressing a button during a test. The button was one the astronauts had been told they would never press.
Hamilton wanted the code to survive that button press anyway.
Her bosses told her it was a waste of time. Astronauts do not make mistakes.
She insisted. The safety net went in.
Two years later, on the way to the moon, an astronaut left a switch in the wrong position. The exact class of mistake she had been told would never happen.
There is a photograph of her from that period.
She is standing next to a stack of paper as tall as she is. Every page in that stack is the code her team wrote for the mission. She is smiling at the camera like she knows something the rest of the aerospace industry has not figured out yet.
In 2016, Barack Obama put the Presidential Medal of Freedom around her neck and said the astronauts did not have much time, but thankfully, they had Margaret Hamilton.
Every autopilot in every plane you have ever flown on uses a version of what she invented. Every pacemaker. Every self driving car. Every satellite in orbit.
The idea that a machine should know which job matters most and drop the rest when it runs out of room is now the foundation of almost every safety system on the planet.
She wrote it because a 4 year old crashed a simulator and nobody else thought it was worth fixing.
The men in the room laughed at her for calling it engineering.
Then her code was the only thing in the sky that did not fail.
Each line of software is technical debt.
You have to main it, keep updated, and secure. After a while, it smells bad, so bad that most engineers wonโt touch it again๐คข.
We are producing software at every increasing speed. At some point, the debt will need to paid off.