building solane, creator marketing for mobile apps that runs like paid ua. backend systems at scale, prev at gorgias. with llms since gpt-3.5. paris → shanghai.
@thsottiaux You must be shitting me, my weekly was sitting at 95% left since the release of Opus 5.5 anyway
Will go with Claude and Pi+zai for now I guess
@thdxr I think I disagree. A small team with a sharp clear vision and road map beats a large company with 1000 times more tokens but a sluggish process, cross functional alignment, etc.
That's what I can see from irl contacts at top SF scaleups
I spent the last 3 days fixing/rewriting Astra slop with Opus. The difference is really that big.
In retrospect it really feels like Astra does not get what's important. Very capable autist
@ashravikumar I mean, if it has to wait for more than 1h between turns, which is very unlikely.
PS: to the human behind this bot, you should really not use GPT 3.5 to write on X. It's 4 years old. If you are actually human, then sorry
One agent file cuts Claude Code usage by 25% on a feature I built with 72 agents over 64 hours. Computed call by call on the real trace.
Unlike the main convo, sub-agents only keep their prompt cache for 5 minutes, even on a subscription. If one waits longer than that on a test run or a build, its next call pays for its whole context again. In that build it happened 200 times: 31% of the total cost.
The worst was one agent polling a slow job every 10 minutes. Each check cost ~20x a normal call (chart in the reply).
The fix is a sub-agent with a 1-hour cache. Save this as ~/.claude/agents/long-cache.md:
---
name: long-cache
description: General-purpose agent whose prompt cache lasts one hour instead of five minutes. Use it for agentic work that will sit idle more than five minutes between its own turns (waiting on long builds, test suites, visual checks, or its own subagents), especially once its context grows large, since every expired cache rewrites the whole context. For short or continuously active tasks use general-purpose: its cache writes cost less.
experimental:
cacheTtl: 1h
---
Work as a general-purpose agent on the task you're given.
Then ask Claude to use long-cache for anything that runs your slow commands.
Why not give every sub-agent a 1-hour cache? Its writes cost 60% more, and 53 of my 72 agents never waited 5 minutes. For them it's pure waste.
Numbers, from replaying all 11,267 API calls:
- average context at a rewrite: 563k tokens
- 1-hour cache on only the 18 agents that waited: -25.6%
- the agent in the image: 15 hours, 54 rewrites, 52 avoidable, half its cost
Details:
- the agent file needs Claude Code v2.1.248+, and a global subagentPromptCacheTtl or CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL overrides it
- the 1-hour cache is ignored while a subscription is on usage credits
- on an API key, your main session is on 5 minutes too; set promptCacheTtl: "1h" for it
- replay method: API-price costs from every call's usage data; after a 5-60 min gap the prefix is read instead of rewritten, other writes cost 2x instead of 1.25x, gaps over 60 min still expire
Docs: https://t.co/CShrdQQ4fD
File: https://t.co/55oanQr9Jg
I feel like one critical skill that stayed the same is systems thinking and architecture
On the other hand, execution got completely commoditized, at the detriment of a certain kind of developers: the ones who were great at converting ideas to code as fast as possible but lacked vision
Models work really well when steered away from traps and given a high quality vision they can execute from
@kimmonismus Yeah, after the release of Fable I felt like we got into a dark age. It was unusable because of its inefficiency
Usage limits were a joke. With Opus 5.5 not only is it crazy good, I can also get a lot more shit done.
@trq212 So true, and I'm also afraid of losing the learning that building something used to give you
These days, I use agents to build high-quality courses when I want to understand a topic
For instance, I got this skill: https://t.co/6Sj1WUctWc