The @karpathy's "car wash" prompt revealing the jaggedness of AI models still works and I found a simple solution: Let another agent watch the given conversation and point at absurd checkpoints.
What if there was a constant watcher correcting these occurences even in less obvious situations during the day?
Might not be obvious but biggest threat to Anthropic is Claude Code itself. People can use its capabilities to build better models, better harnesses and eventually leave.