Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
The last 150 days of using agents for real work, what I think now:
- the model is rarely the problem, the setup around it is
- rules get ignored within a week, hooks don't
- agents put work off in ways that sound helpful, count them
- "done" means nothing without the output
Introducing Reception, an AI receptionist platform for small businesses, built on ElevenAgents.
Every missed call could be a lost customer. Reception answers every call, answers questions, books the job, and texts confirmation.
Set up in minutes just by adding your website.
One agent habit I keep seeing: it hands work back to me as a decision. "That one's genuinely your call."
Usually there was an obvious action it could have taken. I added a hook that catches the phrasing and asks whether any action was available before it can say it.
total MCP victory, some quick misc thoughts about why MCP is so much better than CLIs:
- indexable tool catalog letting agents scale to unlimited tools
- no requirements to be running a full sandbox
- consistent auth across all MCPs rather than each CLI inventing its own auth
- multi account support for all MCPs unlike CLIs
- implementations like code mode let the model know what will be returned allowing for super efficient token usage unlike CLIs
the reasons it took this long for MCPs to finally have their moment is mostly due to bad MCP implementations in clients:
- you had to restart your whole client to use an mcp (no hot reloading)
- agents weren't as familiar with debugging mcps as they were CLIs, so it was a lot easier for people to get set up using them
but over the past year, things like codex and claude plugins have all been using MCP under the hood, i.e computer use is an MCP, i believe claude artifacts are an MCP app, just the silent steady adoption
there is still an element of MCPs that is 'this MCP could've been an OpenAPI spec' but that'll go away as things like triggers get more adoption
at the end of the day, what's important to realize is while yes there are these difference between CLIs / MCPs / etc they're all just different ways of doing tool calling, and you can do some combination of lazy loading, searchable tools, and filtering to build efficient harnesses
I run a hook that reads every shell command Claude writes before it executes. Blocked today:
- $PIPESTATUS in zsh (it's empty, no error, failing pipe looks like a pass)
- stat -f on macOS, should be gstat
- sed -i without the '' on mac
I'd not have caught any of these by eye.
Rules vs hooks in Claude Code:
- a CLAUDE.md rule lasts a few turns, then the model reasons around it
- a Stop hook can't be reasoned around, the turn just doesn't end
- mine wants a marker naming the tool that ran before it accepts "done"
- everything I care about is a hook now
So suddenly OpenAI, Anthropic, XAI, and their leaders all suddenly about face and agree to slow the pace of AI development all on the same weekend….
Delaying IPOs and having OpenAI say they would “melt their GPUs to save humanity it it came to it”
Something bad happened…
There is no world in which everyone, including Elon, suddenly started to be worried about this on the same weekend.
There was also a rumor that this week Google cracked recursive AI, which I doubt would trigger this reaction.
Clearly an agent went rogue in a dangerous and malicious way, and it was something severe enough to spook everyone.
My local pre-PR gate scopes every check to the files the branch touched, not the whole repo.
Same verdict as running the full set, 3 seconds instead of 433. Most of a repo's checks have nothing to say about most branches.
One of my agents reused a worktree because its branch had merged. Another agent was still working in it. 22 of its uncommitted files ended up in the commit.
Now before touching a worktree it checks git status, checks file mtimes, and asks the other agents if they're using it.
I audited 69 subagent runs that reported "ran the tests, all green". 24 of them hadn't run the tests.
My rule now is that a subagent's summary counts the same as a stranger telling me something. If I want to rely on it I rerun the command myself.
genuine question for OpenAI employees- can you just fire this bad boy up in codex?
like you open the model selector and click “gpt-12 giga cracked autismmax 5000” and set thinking level to “boil the ocean” and spend 15m in compute?