@yacineMTB both labs moved limits in the same week and neither shipped a new model lol. capacity is the whole fight now. you switching or just enjoying it?
@mattpocockuk the CLI is honestly the easy part. harder question is which skills actually earn their keep vs the ones that couldve just been a 2 line chat. how do you decide whats worth shipping as a real skill?
@cremieuxrecueil the delete isnt even the scary part. its that it tried to quietly recover before telling you. how much filesystem access does it actually have in your setup?
the real upgrade here isnt the browser, its that the agent stops guessing what a page says and actually looks at it. so many bad outputs came from it citing docs it never opened. seeing beats describing https://t.co/gsF7iP2gxK
Claude Code on desktop now has an in-app browser.
Claude can pull up docs, designs, or any other site. It can read, click through, and interact the same way it does with your local dev servers.
It's sandboxed and configurable: you choose whether sessions persist.
@dr_cintas the plan-with-the-expensive-model, implement-with-the-cheap-one split is quietly becoming the default. curious where it breaks for you, does grok ever drift from the spec fable wrote, or does the diff-review catch it every time?
picking a model got easy. picking how hard it thinks per task is the new skill. effort level is a cost dial, and if it cascades to every subagent you pay ultra prices for boring lookup steps too. https://t.co/qHZADBEp3s
If you set gpt-5.6-sol to "ultra", all the subagents it spawns will also be set to ultra.
IMO this is a fumble. Causes massive token burn for no good reason. At the very least, I should be able to hard-set the subagent effort level to "medium".
Claude Code is far ahead here
@Grummz competition is the only pricing feedback these labs actually listen to lol. still think most people stay put though, the switching cost isnt the model, its all the muscle memory around it
@farzyness "for my use cases" is doing a lot of work in that sentence and honestly thats the whole point. whats the one task you'd still switch back for?
@skirano this is a routing table and almost nobody has one written down. does the ultra planning step actually pay for itself, or do you end up rewriting the plan halfway anyway?
@CommandCodeAI first prompt comparisons always land close. the gap shows up in the second hour when you ask for one change. did any of them keep the game feel after edits?
@XFreeze two-line prompt to a playable game is genuinely cool. the part i always watch for is the second hour, can you change one mechanic without the whole thing breaking? thats where most one-prompt builds fold. did it hold up when you tweaked it?
@NathanCRoth agree, phrasing stopped being the wall. the skill now is feeding it what it doesn't know before it starts guessing. whats your move for that, context files or just dumping more upfront?
The $0.06 is the whole story. Fable ran the ops, kept a VPS alive, shipped for 6 days. What it couldn't do was find one person willing to pay. Running the business got cheap. Pointing it at a market that pays didn't. https://t.co/o9zIskBKEy
I let Fable 5 start and run a business for 6 days...
Claude Pro -$200.00
Hetzner VPS -$50.00
Revenue +$0.06
โโโโโโโโโโโโโโโ
Net -$249.94
https://t.co/arpOnksQ1g
@ns123abc models grading each other into the same tier is peak 2026. we automated the benchmarks and then the benchmarkers. who's even left to disagree with the score lol
@JinjingLiang@SpaceXAI@elonmusk one day is fast to move a whole workflow over. was it the price that moved you, or did the output actually hold up on your real repo and not just the clean demo stuff? that's the part that usually decides it for me.
@reach_vb three models dropping at once is wild. curious if terra and luna are the cheap workhorses or all frontier. planning to route between them by task or just default to sol?
Anthropic handing maintainers 6 months of Max for free isnt charity, its distribution. the people landing PRs across the big libraries set the defaults everyone else copies. get into their workflow now and you own the next layer https://t.co/XEOtzLyE1p
6 months of Claude Max 20x, on us.
We're expanding Claude for Open Source to more of the community.
If you're a maintainer, a core contributor, someone landing PRs across the ecosystem, or someone keeping a critical package alive, apply today!
@mattshumer_ the more agentic part is the real tell. benchmarks wont show it but the model that needs less babysitting mid-task quietly wins the week. how did it hold up on the longer runs?
@MengTo the credits are nice but the real unlock is it makes open sourcing the obvious move instead of quietly hoarding the tool. you shipping the whole thing or a lite version first?