Ran my own tool on my own machine last night and found 463 errors I'd never looked at.
Built walkaround back in June. It reads Claude Code's local transcripts and reports what the agent actually did to your repo. 40 sessions on this laptop. 113 errors in the headline numbers, the ones I read every time. Another 463 sitting inside the subagent rollup.
I'm the one who put that rollup in its own section, so sessions with and without subagents would stay comparable. Then I stopped opening it.
Fixing that next: one line per subagent instead of one for all of them. Codex after, if someone can tell me what those transcripts look like.
13 developers have tried Walkaround thi week 🚀
I’m really proud of that.
It’s not just a number; it represents progress: a milestone of success to build upon as I continue creating value.
One week after launching here on X, it has garnered 13 genuine downloads and analyzed Claude Code sessions, uncovering errors, metrics, and hidden processes.
I am convinced that this small success can evolve into a tool capable of providing comprehensive analytics.
As previously mentioned, sub-agent analysis is coming soon.
Still working.
I partly agree with you. I think we are increasingly moving in a direction where we will no longer have "power and control" over the actual coding, but rather over making the software infrastructure secure, scalable, and manageable by autonomous agents.
I find the testing phase more important now than ever.
I don't know if I would currently entrust a project 1000% to AI for total management without supervision; instead, I would likely put more effort into testing, especially regarding security and transparency.
@marclou@loaibassam@marckohlbrugge An excellent trade-off to avoid adding too much friction to the offer proposal and secure receipt processes.
I’ll keep this in mind for a project I’m working on!
A few weeks ago, I tasked Claude Code with prospecting for potential clients; it deployed around a hundred agents across X, Reddit, LinkedIn, and various forums.
At one point, however, several of them stalled.
I investigated, expecting a technical glitch.
It turned out the model’s safety protocols had kicked in, complete with an explanation: one agent was compiling dossiers on real people based on their weaknesses, while another was scraping personal data by bypassing site blocks.
I was looking for clients, not targets.
From the outside, though, it’s impossible to tell the difference: a file listing a name, product, estimated revenue, and alleged technical vulnerabilities looks exactly like the preparatory phase of an attack.
Models often go off-track and tend to fear the worst.
I’d call them "pessimistic" and, to be fair, the line between my objective and what Claude feared was very fine.
What do you think about model guardrails?
Do you find them sensible or excessive?
@forgebitz The most fun thing ever is watching "works on my machine" become "worked in my session". Do you ever re-run the same prompt hoping it'll break again so you can watch it?😂
@johnrush AI erased the best part, and I'd narrow it to the writing. Reading and deciding can absorb you the same way, but it's a different job, and nobody spent twenty years falling for that one. Haven't hit sunrise on a code review yet
What moved it for me was banning specific words instead of describing a style. "Write clearly" gets ignored, a list of banned words and no em dashes doesn't. Tying every sentence to an actual change in the diff kills most of the waffle. The bullet list reflex I still haven't beaten.
Yeah, and it goes wrong on the way back up too. What returns is a rollup the subagent wrote about itself, and when I parsed my own session transcripts that's where the numbers drifted most from what actually ran. Does your harness give you the raw subagent output, or only its summary?
@addyosmani Autonomy earned by passing verification loops depends a lot on what the loop reads. An agent's own summary comes out of the same run you're checking, so it can pass for reasons that have nothing to do with the work. How do you decide what a gate has to look at?
9/9
Remember the landing page I keep not taking down?
It's this one: https://t.co/FaYsqxVzVb
Still online, still says "Coming soon".
Something did come. It just wasn't TweeX.
1/9
Back in February I wanted two things: to stop being inconsistent on X, and to make my first euro online.
So I did the most builder thing possible: I built a SaaS about it.
It died four days into beta.
The story is better than the product was. 👇
8/9
Today my X runs on Claude Code: a repo of instructions it reads, drafts it proposes, me approving every single publish.
Zero X API, zero per-seat costs, just the subscription I already had.
What I was trying to sell was a workflow all along.