Just saving this here to document a story and as a self reflection on whether AI is really making me more productive
Yesterday morning I found a way to complete the new HVM approach, that is much faster than before. I spent a few hours writing a spec, and then used Opus to implement. About 3k lines of C code later, everything worked and performance was incredible: 5x faster than HVM4 (stable at ~10x now). So, in one day I had outclassed HVM4. Incredible. I'd never have implemented that so fast manually.
Now, enter today. I want to turn this into a real thing, but I haven't fully read the 3k lines yet. So, how do I trust it? I spent the whole day auditing the code. With AI. Several bugs found, most minor like forgetting to collect() some argument. But then I stumble upon this:
λ{ inl: 1 ; inr: 1 }
This was a test. But wait. This is matching on inl/inr. So the branches should receive the value of the Either. But they were numbers instead. Numbers aren't functions. This makes no sense. So why this is a test?
It then stuck me. The AI completely misunderstood how function arities work. It literally assumed for no good reason that HVM5 was supposed to handle under/over-applied functions. For no good reason. I never wrote that. It never asked either. It just kinda thought "HVM is weird in some aspects, this might be one of them..." - and then it went on to implement a massive system to handle cases that should never happen to begin with. And all of that code is obviously wrong because it should not even exist. It is wrong. It is damage. And it is there.
But it isn't too bad either. I just told Opus that it was wrong. Perhaps not so politely. And it solved it just fine.
But then this begs the question. I spent ~20 hours in this file, and it is STILL not done. I went from 0 to 95% in the first 5 hours. Yet, 15 hours later, it is still not 100%. I suppose that is the real effect of using AI. If I had just written the C file manually in the last two days, would I not be further than where I am *right now*?
Surely, the first version would have taken much longer to drop. But when I'd finish writing all that code, there would be zero, literally zero retarded shit. And, just today, I caught 5 or 6 retarded shit. And the worst part is: I don't know what the number of retarded shit left is, but I'm afraid it is >0.
So if I have to read it all, review it all to ensure there is no retarded shit... what did I achieve by using AI, other than that dopamine anticipation?
One of the biggest problems with using LLMs as a google replacement for programming, is that getting zero relevant results on google used to be a signal that you had the wrong idea about the root cause. Whereas LLMs will happily indulge any terrible idea you suggest.
LLM psychosis scales with your distance from the code. As a result it tends especially to afflict non-coding managers, PMs, and execs. It’s also a self reinforcing loop. As the code becomes an object of disgust (unreadable pile of vibecoded shit) you are forced to distance yourself further from it and your only interactions with the code are mediated by model.
I’ve gone through the whole arc of
1. Omg AI can write code
2. Meh it makes mistakes
3. Omg Opus is AGI
4. No it still makes slop
5. Yes but everyone is still gonna use it like crazy
And my view is these things need finer context.
If Codex was told this is high performance pathway and defensive guards have already been applied before code enters this flow, it would have not written it like this.
Need to manage context more finely and you can still deslop these things.
- how frequently to iterate
- how long a loop to run on auto pilot
- how often you need to interject and deslopify with these finer contexts
- up to which depth of the tree you need to write further AGENTS.md or rules and guidelines
All of this is still a very “it depends” level and I am seeing good engineers learning the ropes of effectively doing all that.
Napoleon didn't tell his generals "you've used enough cannons for today, come back tomorrow"
So why is Claude rationing my sessions when I still have weekly ammo left
In an industry built on determinism, I feel we might be underestimating the work we all will need to do with LLMs exactly because they are nondeterministic.
But for so much of automation/workflows, determinism (aka "make sure it doesn't make a mistake") is a baseline expectation
I think we’re speed running understanding that when a company controls the model and the harness, and they are both closed: they not only CAN pull stuff like this, but WILL do so.
I expect a renewed interest in open models, open source harnesses + self hosted models. This stuff is getting really disruptive and is just not acceptable as a paying customer!
These are genuinely interesting times to be building software and infact I’ve not had so much fun building new software in a long long time. But these are also times where it pays to be clear-eyed about what you’re trading away.
https://t.co/ULpNTy2LWE
https://t.co/BGfGZBMGe8
this isn't a huge deal but this is really the flavor of our times
everything is just sloppy. everything has 20% margin of error. nothing has precision
just automate it. just select all. just make ai figure it out
i do it too and it all adds up to a gross feeling world
@BlueDartCares Your local staff are a nightmare! They constantly lie, marking my bank docs 'undelivered' and returning the shipment without delivery. NO working helpline or way to contact the delivery boy. Absolutely pathetic service.
My neighbours stay in their own bungalow valued at Rs. 20 crore and yet drive a Nano. They are 65+ and the lady of the house finds it convenient for driving to nearby market etc. When ego isn't a factor there is not much we need.
Life’s hard. You meet a girl, fall in love, get married. Her dad wants you to work 70 hours a week. You can’t work that hard, you just wanna chill and run England.