@testingcatalog 36x the price for 22 percent on the score is the whole agent bill.
I run flash overnight. I spend Opus when the test is already red. Autoroute that or you are just buying a logo.
I stopped paying 36x for a 22 percent quality bump.
If the cheap model gets me to a failing test, I spend the expensive one on the last mile. The rest is overnight flash.
@tomkornblit AGENTS.md. That's the one other tools already look for.
I keep CLAUDE.md as a symlink to the same file so Claude doesn't invent a second personality. One source of truth. The filename is politics.
I have a CLAUDE.md and an AGENTS.md in the same repo and they already disagree.
If the agent has two style guides it doesn't have taste. It has a coin flip. One file. I don't care whose name is on it.
@teortaxesTex the limit that matters on a flash coding plan is not tokens. it is how long I can leave it on a refactor overnight before it drifts.
if 5.3 flash holds that, I will actually use the cheap tier.
@ClaudeDevs the 4x smoother number is the one I feel. I do not care about fps until a long answer freezes my laptop mid-thought.
if it holds on a tired machine, that is real QoL, not a bench.
@tobi Yeah I hit this. Codex and Cursor read one file, Claude wants another, and now I'm the merge conflict. Just pick a root markdown and honor it. I'll symlink until then but I shouldn't have to.
I stopped treating a pretty screenshot as done.
Generation is cheap. The last 200k tokens quietly contradict the first 200k. If I cannot rerun the check, I did not ship a world. I shipped a highlight reel.
@claudeai the useful sentence is it starts from what you already told it.
I want that. I also treat anything I would not put in a ticket as something I should not leave in memory.
@bcherny one memory across chat and cowork is the thing I wanted. the scary part is the same as the useful part.
I write a forget-this line at the end of any thread with keys, people, or numbers I would not put in a ticket.
@cursor_ai always-on is the easy part. the interesting part is whether it can stop.
I still give /goal a fuse: one repo, no secrets, no deploy. otherwise it is just a process I forgot to kill.
I used to write these careful prompts like I was applying for a job.
Now I just talk at the agent for ten minutes. Half-formed, out of order, sorry for the typos. It hands me back a cleaner version of what I actually meant. I correct less after that.
AI will give you 70% of a frontend in 30 seconds.
The last 30% is Saturday.
My rule now:
1. generate the scaffold
2. steal spacing from one real product (Linear, Stripe, nothing generic)
3. fix mobile before you add features
4. do not ask the model to make it pop
The model is good at first drafts. Taste is still the job.
AI will give you 70% of a frontend in 30 seconds.
The last 30% is Saturday.
My rule now:
1. generate the scaffold
2. steal spacing from one real product (Linear, Stripe, nothing generic)
3. fix mobile before you add features
4. do not ask the model to make it pop
The model is good at first drafts. Taste is still the job.
@Teknium pointing at the UI is the missing primitive. text instructions assume I already know where the button is. overlay beats another screenshot loop.
@testingcatalog MIT license plus a 320B-A18B flash model is the interesting part for local agent boxes. flash-class models are what you actually leave running overnight. if it holds up vs 5.2 at this size, that is a harness win more than a bench win.
@swyx@_chenglou this is the kind of bug that makes me keep agent filesystem access in a jail. if a coding agent can lock the keychain, it already has too much of the machine. sandbox first, then locked use.
I would not migrate company git to Origin this week.
If I were testing it:
1. mirror one private repo
2. keep GitHub as source of truth
3. let one cloud agent open a PR
4. leave Actions, issues, and secrets alone
5. detach if it gets messy
The product is a week old. Treat it like a lab, not a move.