Opus 5 is a great model for coding, data analysis, design, biology, knowledge work.
More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.
And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon.
https://t.co/Tc7z2FqJhQ
@bcherny thanks so much for telling how Claude works ..
Actually everything else especially loops feel so dumb when you really use the power of Claude !!
https://t.co/8sr8mwnZmi
This whole loops thing sound so unbelievably stupid!!!
Sorry guys, but hey if you are smart you hire a professional!
You don’t tell an idiot for forever!!
This is the EXACT architecture OpenAI uses to build AI agents.
They just dropped a 34-page guide, I compressed it into one page.
6 stages, one loop, everything you actually need.
Read it, then go to the step by step guide on building LOOPS for your agents below.
A senior Anthropic engineer just dropped 11-page PDF on "Loop Engineering" for agentic systems.
The shift: you stop prompting the agent. You build the system that prompts it instead.
Schedule → Discover → Build → Verify → Repeat
Every loop runs one turn, five moves:
• Discovery: it finds its own work - failing CI, open issues, recent commits - instead of being handed a list.
• Handoff: each task gets an isolated git worktree so parallel agents don't collide.
• Verification: a second agent, told to assume the code is broken, reviews the first. The "thing that can say no."
• Persistence: results get written to disk, never left in a context window that gets flushed.
• Scheduling: an automation wakes it on a timer. That's what makes it a loop.
The key insight: an agent grading its own work always praises it.
This 11-page PDF changed how I'm building agentic systems today.
Read it now, then explore the article below.
A senior Google engineer dropped a 424-page doc on agentic design patterns.
424 pages.
Most engineers bookmarked it and never opened it again.
I read the whole thing.
Here are the 15 patterns that actually matter — explained in plain English, with exactly when to use each one ↓
R.I.P. paying full Opus prices for every single AI task.
A properly routed open-source Claude stack can replace $200+ a month in frontier model spend.
It is not as easy as just swapping the model name and hoping for the same output.
But if you start today, you can have GLM 5.2 wired into Claude Code, a local model running on your machine with zero token cost, and your first autonomous loop built, verified, and running unsupervised by end of this week.
I usually charge $99 for access to this playbook but today, it's free.
Like this post + comment 'STACK' and I'll DM you the full guide for free.
The guide covers three things.
How to set up local models on your hardware in 15 minutes using Ollama, which model runs best at your RAM level, and the decision engine that tells you which of your tasks belong on local, which go to a cheap API like GLM 5.2 at $1.40 per million tokens, and which 20% actually justify Opus.
How to wire GLM 5.2 into Claude Code in under 5 minutes by editing one JSON config file so the same harness, skills, and workflows you already have run on a 5x cheaper engine for 80% of your tasks.
How to stop prompting and start building loops. The 4-condition test that tells you which tasks are ready to loop, the four blocks every loop needs, and the copy-paste prompt that builds your first loop orchestration skill with training mode, memory, and a verification step included.
(Must be following, or I can't message.)
Taking this down in 48 hours.
Anthropic engineer:
"You're not supposed to watch Claude Code work. You're supposed to wake up and review what it shipped"
in 22 minutes she builds the full workflow from scratch, step by step
a routine that ships work while you sleep
most developers have never seen Claude Code run without them watching
watch the session, then save the guide below