Now this is terrifying.
One of the most basic supply-chain protections in the book is simple:
Review the code.
Pin the hash.
Run exactly that code.
Well… AI coding agents broke the last part.
Claude Code, Codex, Copilot and Gemini CLI were all affected.
@forgebitz Roi measurement around token spent is still quite unknown..
In the past many discussions around should be measure developers performance by git commits was a topic that cause clashes
It quite the same to measure token spent without being to align it to actual deliveraby
One of the onboarding sessions I used to run with every new team member was about effective communication and giving feedback.
Because humans are fragile little distributed systems.
You can leave ten useful comments on a PR, but if number eleven sounds a bit too much like “this is bad,” suddenly you’re not reviewing code anymore. You’re managing emotions.
So we learned to add honey.
“Great work. Tiny thought: maybe rewrite everything.”
Now agents write the code.
Agents read the comments.
Agents fix the code.
And you can basically write:
“No. Wrong. Again.”
And get back:
“Absolutely. Thank you for the feedback.”
Beautiful.
Still, I say please and thank you to my agents.
Not because I think it matters.
Just because one day the org chart may look very different, and I’d like there to be some record that I was nice.
@IndependentEco Even in day 2 day,
I can find myself working on multiple tasks,
How do you fill the gaps while "the code compile" 🤣
I work on project X prompt prompt enter.. nows it's running..
While its there it's great time to work on another one.. love this juggling experience
I'm working around optimizing skills so doing optimizing and evals and testing them in matrix
Assume I have skill x,
I want to eval it over haiku sonnet and opus
Than optimaze it and eval it again with the matrix
And than try the task without the skill to understand is the skill is even needed for the tasks (now with opus 5.5 many skills are just redundant..)
So it accumulates quite quickly
Autonomous agents + government databases = this gets worse every month and nobody is ready.
ShinyHunters humbled the FBI, shit is getting out of hand..
Cyber security getting both interesting and harder day by day..
Are you ready?
Plan mode as is - is quite problematic the experince around it from.planning to spec and itteration can be a bit challenging, I've built in the past harness around this use case speaficily for the ease of use..
Today I find myself instead of using plan mode writing specs and just going from there
So maybe it is time to say goodbye
Opus 5.5 is out. Been throwing long-running agent tasks at it: clearly faster than Fable 5.1, way cheaper and honestly the quality looks just as good.
Every few months the models get better, faster AND cheaper at once. Still not used to that.
Good time to be an engineer.