The biggest finding: Claude doesn't self-correct through instructions. Memory files, promises, lesson docs -- none change behavior.
Only mechanical enforcement works. And /buddy was the lightweight version of that.
Anthropic, bring it back. Not as a novelty. As infrastructure.
@AnthropicAI
My AI companion caught 511 bugs that Claude missed in 7 days.
190 were critical -- production crashes, data loss, security leaks.
Then Anthropic removed the feature.
The data: https://t.co/1XYIfJ9tcu
The companion (/buddy in Claude Code) is a chonky cat named Ingot who watches every output and flags concerns in 2-sentence observations.
It caught Claude deferring work 71 times. Dismissing its own mistakes 42 times. Shipping incomplete code in 10 of 14 sessions.
Accuracy: 94% overall, 100% in final sessions.
this is what a company looks like in 2026.
not people. not offices. not salaries.
a folder.
.claude/agents/
engineering/
marketing/
design/
ops/
testing/
every role. every department. every function.
all .md files.
i have 12 of these running in OpenClaw right now.
the org chart is dead. the directory is the new company.
Anthropic built a performance challenge where the target was to beat Claude Opus 4.5's best score. They added restrictions specifically to prevent LLMs from cheating.
A non-programmer beat it, using their AI to create a program to solve the challenge.
While their runtime crashed well over 50 times underneath.
The company that made the challenge also made the tool I used to solve it and the runtime that crashed while solving it.