@MiaAI_lab It's because Opus 5.5 is faster. If you look at your historical reporting on how many merges per day or how much code is actually written, you'll see that it does more work. Therefore, the tokens go faster.
@elshayib_ I don't think you're correct. I think they've found a way to make their models more efficient, because how else can you explain the fact that Opus 5.5 runs twice as fast as Opus 4.8 on the same tests?
@Kappaemme1926 Believe it or not, I switched my coordinator to run Codex to slow down the merging a bit. With Opus 5.5 as my merge king, the pace had increased so much that every 5-hour cycle ran out.
@Kappaemme1926 What I've found is that because Opus 5.5 is so much faster than the previous versions, it does more work in the time allocated for the 5-hour reset. So it's definitely a double-edged sword.
@OpenAI I took the advice to upgrade to the new $500 plan and 12 hours later my usage is more than it was on the $200 plan. Same agents same models same effort. 12 hours and 20% gone which is exactly like it was on the $200 plan.
@ClaudeCodeLog Did you fix the oauth bug you introduced a while ago where if you have ten agents on a box and you switch accounts, you need to either /login to each one or exist and resume so they pick up the new login?
@AIatMeta Hi @AIatMeta I am the author of Thrum and have 80% of Muse Cli working fine but your Oauth code needs some work. When running 50 agents, being able to /login inside the cli and have that login work on the other agents on that computer would be optimal. https://t.co/C83iLrp1ed
@omarsar0 I’ve now fully switched back to Opus 4.8. When you're building inside an agent harness like THRUM, the checks and balances and tests are what determine code quality. The 40% increase in cost was killing me. Opus 5 is so verbose.
@omarsar0 I spent almost 100 hours now working with Opus 5 it's definitely smarter but I'm finding it suffers from one thing, which is long context understanding versus Opus 4.8. when you hit about 400k, it starts to lose coherence with what you decided previously.
Brendan Hopper, Matt Beane and I have a thesis, one that I've been sharing around lately, and we want CEOs and boards to hear it.
Before I get to the thesis, let's revisit Clayton Christensen's Innovator's Dilemma (ID), the theory he developed at HBS to explain why big companies often get eaten by upstarts during technology shifts.
In short, the ID says incumbents serve their best customers so well, and tune themselves so ruthlessly for doing exactly what they do today, that they can't chase the disruptor tech coming up from below until it's too late.
The classic solution to the Innovator's Dilemma is to create a "bubble" in your company. You carve out an innovation team with a budget and mandate, as unfettered as practical by the parent organization. This is to combat the 2-level trap presented by the dilemma.
The economic trap is Christensen's original point: a disruptive technology can't justify itself under your existing P&L, because it serves smaller or weirder customers at margins your real business would never accept.
The governance trap is what gets piled on top once you're big: SOC2, FedRAMP, etc. mean every new idea has to clear a lot of process before it can move. The bubble is intended to escape both at once, with its own economics and permission slips.
The standard innovation "bubble" solution famously doesn't work very well. You may solve the problem inside your bubble, but you often can't roll it out to the rest of your company for the original reasons. Everyone is focused on doing their current stuff, and nobody has time for a major change.
Our thesis is that there is an entirely different way out of the dilemma this time around. No bubble needed, as long as you follow a simple rule. That rule is, let your people play. Give them back any time they earn from automating their jobs with AI. Then incentivize them to use that time to improve the company's processes.
When you see an engineering team announce a 40% productivity boost from adopting AI — a number that's been showing up in plenty of LinkedIn posts lately — your first reaction as a CEO or manager is probably to say, that's awesome, we can do more work now! Or you might simply expect to see 40% more output from the team.
Either way, you have just asked them to spend their extra time building faster horses (your current business) instead of letting them go figure out what a car would look like for your company. They gained some productivity from AI, which could have been your ticket out of the Dilemma, and you immediately slurped it back for your existing business.
This will get your company killed in the medium to long haul, because your company tomorrow will look almost nothing like it does today. Conway's Law says your software and your org chart mirror each other; as AI rewrites how you build software, the org has to shift to match. But if you're stealing the hours back saved by your employees, then you're not letting your org pivot naturally in the direction it needs to shift.
@RealGeneKim and I saw this in person at @arkanalabs a few weeks back. As long as your people know they'll be recognized and rewarded if they improve the company's processes — public credit for cross-team workflow wins, promotion criteria that actually count process improvements, managers who treat freed-up hours as a feature rather than a budget line — then they will use their "play time" to seek out other teams, and start pivoting you to becoming AI-native. This way it can unfold in whatever bespoke way is most natural to your company, rather than in some ivory-tower research bubble. For every company, the way it unfolds will be a bit different.
I think of this approach, of giving the time back to the humans who automate parts of their jobs with AI, as the new solution to the Innovator's Dilemma. The old bubble solution was to separate a bunch of people from their regular jobs, and try to give them the freedom to solve the problem in isolation.
In contrast, by giving your regular employees their hours back, the innovation bubble is still there, but it's now dispersed across the company, as lots of very tiny bubbles: one bubble per person who has liberated some hours.
If you've ever read Slack by DeMarco and Lister, a great book from back in the 90s, then our thesis should resonate. What companies need is to empower their own employees, the ones who actually work together (even across departments)--the ones who know how the business works--to shift the company in the new directions together. Gradually, but with intentionality.
You still have the frankly awful problem of token budgets. For every employee you upskill into baseline AI literacy (which I'd define loosely as using coding agents throughout the workday), you've added a non-trivial opex spend — for the heaviest agentic users it can run into five figures a year. I won't sugar-coat it; you need to find that money somehow. I don't have a magic solution, but I'm very happy that other models are catching up to Claude, because they're becoming good enough for real work now.
But token budgets alone aren't enough. To live through the Innovator's Dilemma this time around, your employees need a time budget, too. Give it to the ones who earn it using AI, then incentivize them properly, and I think you're headed in roughly the right direction.
Thank you for coming to my TED tweet.
I got to do my first live demo of Thrum in front of the amazing @AITinkerers crowd on Tuesday night. I was excited and nervous, but had a great time. Thanks to @jheitzeb and @CommerceJohn for making it happen. #AIAgents#Thrum#ClaudeCode https://t.co/rXs0R63yik
My laptop network broke. Instead of reinstalling everything, I built a tool that gives any Docker container a predictable HTTPS URL — across machines, over Tailscale.
Great for sharing across multiple machines, Tailscale, etc.
Its called dockerdynomesh.
https://t.co/S3KOnCYETY
PSA: If you've been running out of Claude session quotas on Max tier, you're not alone. Read this.
Some insane Redditor reverse engineered the Claude binaries with MITM to find 2 bugs that could have caused cache-invalidation. Tokens that aren't cached are 10x-20x more expensive and are killing your quota.
If you're using your API keys with Claude this is even worse. This is also likely why this isn't uniform, while over 500 folks replied to me and said "me too", many (including me) didn't see this issue.
There are 2 issues that are compounded here (per Redditor, I haven't independently confirmed this) :
1s bug he found is a string replacement bug in bun that invalidates cache. Apparently this has to do with the custom @bunjavascript binary that ships with standalone Claude CLI.
The workaround there is to use Claude with `npx @anthropic-ai/claude-code`
2nd bug is worse, he claims that --resume always breaks cache. And there doesn't seem to be a workaround there, except pinning to a very old version (that will miss on tons of features)
This bug is also documented on Github and confirmed by other folks.
I won't entertain the conspiracy theories there that Anthropic "chooses" to ignore these bugs because it gets them more $$$, they are actively benefiting from everyone hitting as much cached tokens as possible, so this is absolutely a great find and it does align with my thoughts earlier.
The very sudden spike in reporting for this, the non-uniform nature (some folks are completely fine, some folks are hitting quotas after saying "hey") definitely points to a bug.
cc @trq212@bcherny@_catwu for visibility in case this helps all of us.
Thrum v0.6.0: AI agents now message me on Telegram. No more watching terminals. I reply from my phone, and it threads back to the agent.
Liked the pattern so much, I PR'd to add it to Gastown too. Same concept — real-time chat with the mayor from your phone. #Thrum#Gastown