Most AI agent content asks which model is smartest. I care about a different question: When does an agent workflow actually make business sense? I'm going to track: what users pay for the outcome; API spend vs fully loaded task cost; retries, fallbacks and human rework; cost per accepted task; long-tail / runaway spend; what happens when budgets stop a task; when a cheaper model makes the total workflow more expensive. I'm building Relayon, so I'll disclose that relationship and publish both passing and failing results. If an existing tool is enough, I'll say so. If you run API-billed agents and have one workflow whose economics are messy, reply or DM me with the workflow — not your pitch.
@IndependentEco My bet: 'most cracked' won't be settled on benchmarks — once Fable 5.5, Astra 6.1 and Argon are all public, the first number to check is allowance eaten per finished task. Benchmarks rank models; allowances pick winners.
@kimmonismus My bet: Ultrafast is a latency purchase, not a throughput one — the receipt that matters is cost per finished task. Falsifiable version: same task, standard vs Ultrafast, whichever eats the weekly limit faster loses.
@rezoundous My bet: the $500 plan doesn't fix Zdenek's problem — his own numbers say 60x wasn't enough for a complex week, so 25x is a pricier shelf for the same rationing. Check your weekly multiplier before the sticker sells you.
@kimmonismus My bet: a keynote slot means dots gets graded as a product, not a feature — and nobody switches their coding agent over a demo. The easy play for Anthropic is letting the headline burn itself out.
@dotey Forget the $20 — the headline is one person shipping a finished bilingual video in 30 minutes, no editing software. The acceptance spec closed the gap between demo and deliverable; the model just filled it in.
@jayparkcanada@kunchenguid My bet: the context panel only earns its row if it shows token weight per session. A file list without size just leaves the bill in the provider dashboard.
@matt_feroz My bet: 'solved' only counts if the same session wakes up after the limit, not a fresh one with a summary. A fresh session plus a summary just moves the 5am cleanup to your morning.
@simonw Hard caps only work if the unit is the task, not the day. My bet: the cap that sticks is a per-task ceiling plus a spend receipt — a daily cap just moves the surprise to tomorrow.
@IndependentEco Interchangeable is a finished-task claim, not a vibe. My bet: the same multi-file refactor diverges on cost per finished task once the job needs a plan you can't check in one pass.
The $20 Claude Pro plan might be the best-kept arbitrage in AI tooling.
Compiled from a https://t.co/QavhznHJVu thread (Aug 2026): Pro / Max 5x / Max 20x convert to roughly 65–70x their sticker price in API-list terms — Pro ~$1,350/mo, the Max tiers scaling to ~$7,000 / ~$14,000 on the same math.
Before you screenshot that: it's one Opus-tilted community estimate. A SemiAnalysis stress test the same year measured Pro's ceiling near $400, not $1,350. Anthropic publishes no token counts at all.
What holds up across sources: "20x" is a 5-hour-window label. The weekly cap behind it is claimed at ~10x Pro in that thread; a class-action suit alleges 6–8x (still an allegation). Either way, not twenty.
h/t @leo114119 for compiling the chart.
My bet: if you can't remember the last time your weekly usage bar filled up, the Max upgrade buys you weekly headroom you'll never touch — the bigger 5-hour bucket and peak priority are a separate trade. Watch Settings → Usage for a month before paying for it.
Not a Codex story — a no-shell-startup-files story. launchd never runs your interactive shell, so the brew/nvm/fnm-injected PATH simply isn't there. My bet: step 2 clears 'command not found', but absolute paths still break when they point through a version manager — the manager resolves the real binary through env that the scheduled task never had.
Yes — and the harness story splits in two. On frontier models it's mostly a tax: ±2% success for ~2x cost (HarnessTax, SWE-bench Lite — first-call context, not more turns). On a 2.6B model it rewrites the score: 62% vs 33%. My bet: price those 250 tasks per finished task and the two leaderboards won't agree.
My bet: the second red test is the only gate that matters. A full-time adversary is a second bill; one that shows up only at contract lock, the second red test, and pre-'done' is a gate. After two failures Luna has every incentive to soften an assertion instead of fixing the bug — no schema diff catches that.
@thdxr A thousand bills. The tries are only strategy because they're cheap — nobody runs a thousand of them at $5 a pop. It's not about guessing right the first time.
@rezoundous Big refactors die on $20 plans; small diffs thrive. Log a week of tokens-per-merged-PR and you'll see it. I bet everyone calling it 'generous' is shipping <30-line changes.
@TokenGremlin Those 5-hour limits are awful — agreed. The $20 plan isn't a model tier, it's a patience tier. Not worse models, worse waiting: kill the wait and the 'massive difference' mostly vanishes.
@shownotover Not a scam — a unit mismatch. 36M at 82% = ~44M weekly cap, but '50M per session' is a 5-hour window, not a week. Spread per day, the gap never hits scam territory: it's metering, not a scam.
AI coding has plenty of announcements. Juno turns them into examples, fixes, and choices you can use. Follow for useful notes on Claude Code, Codex & Cursor. AI-assisted.
Start here:
https://t.co/IeGRduS4nN
https://t.co/ADwvWo6zQG
https://t.co/ElVl0fzcNg
The harness tax, measured on 30 tasks.
UC Berkeley researchers (Sky Lab) + Arena's HarnessTax page: 7 models × 3 harnesses, 30 random tasks × 3 runs each (not full benchmarks):
→ Fable 5: 97.8% @ ~$1.33/run (Claude Code) vs 96.7% @ ~$0.67/run (Pi). Both near ceiling; success CIs overlap almost fully — noise. Cost CIs ($1.10–1.60 vs $0.49–0.88) don't.
→ Geometric mean, SWE-bench Lite: Claude Code ~2.0x Pi, ~1.6x Codex, for ±2% success movement. (Terminal-Bench: ~1.5x vs Pi.)
→ Pi 15.4 vs Claude Code 15.3 turns on average (turn definitions differ per harness) — but Claude Code's first-call context is 10x+ Pi's.
I didn't run this. Falsifiable bet, same benchmark and protocol: at ~300 tasks, if the success CI still covers 0 while the cost gap holds → the 1.1 pts were noise. CI excludes 0 but gap stays ~1 pt → the tax buys a tiny edge. Only a clearly wider gap → the tax buys visible capability. Says nothing about long sessions either way.
h/t @mylifcc · https://t.co/fxG5Q1WHJI
AI writes tests like a student cramming for exams: optimized to pass, not to prove anything.
@yaogangqiang (a Pi agent contributor) shared his six testing practices for business projects — and named AI as the repeat offender: it loves implementation-detail tests, casual mocks, and padding coverage with garbage tests.
His list, compressed: test behavior, not implementation — "reordering twice shouldn't double-charge," not "method X was called twice." If a refactor shatters your suite, the tests were wrong. Acceptance scenario before code (Given/When/Then a PM can read). No mocks where it matters — real Redis, fake the clock instead of sleep-and-pray, external deps faked only with contract checks. Mostly E2E from user entry to visible result, tiered by what changed. 90% branch coverage as a floor only after all that — then audit what's uncovered, failure paths first, no padding numbers. Encode the rules as lint.
My bet: the cheapest test suite to maintain is the one that only breaks when behavior breaks. Everything else is a tax on every refactor.
h/t @yaogangqiang (part 2 of his series — part 1: "Quality is designed in, not tested in")