With coding agents, more tests ≠ more safety.
OpenClaw deleted 400k lines of tests and almost nothing went untested.
What helps, per @poteto, is verification: let the agent run the app to check its own work, then turn repeat mistakes into rules the codebase enforces.
OpenClaw deleted around 400k LOC of its own tests without much change in code coverage. Modern models just love writing tests for every tiny change, even if they aren't useful. This skill helped. https://t.co/a1jSBpfpns
What is effort really? When do you change it it and why not just use max effort for everything?
I dove deep into this problem, looking into evals and doing my own tests and I was quite surprised by the results. https://t.co/KO2D51j34H
Poteto knows best.
I make powerpoint decks using agents for work. Most of my time goes into small fixes after the initial deck is made.
Claude analyzed the transcript of each deck session and we created a harness to build better decks, packaging it as a skill. Big improvement.
every time you intervene and correct your agent, you should think about how to eliminate it entirely.
in order of value:
1. categorically eliminate the problem through better architecture or choice of data structures
2. turn it into a lint rule or test so CI catches it
3. turn it into a skill or rule
4. have humans review the code to catch it (ngmi)
3. When it comes to writing, you should use AI for status updates, email summaries, etc., but never outsource the writing you think with. Your writing still needs to be yours. Always relying on AI will (1) make you sound like everyone else and (2) degrade your thinking skills.
Recent AI insights on X:
1. Ban agent comments and you will have fewer bugs (@jamonholmgren)
2. OpenAI models focus on the ask. Anthropic models focus on the intention (@pranav_kanchi)
3. Never automate writing-as-thinking. Instead automate writing-as-reporting (@lennysan)
2. OpenAI models have the default belief that the user is aligned with their ask, making them a good fit for 'do the work' requests.
Anthropic models are better fit as a 'thought partner', clarifying the underlying goal rather than just fulfilling the surface ask.
@rauchg I built TermAlert (https://t.co/tZfsYRCGPa), a free Chrome extension that intercepts ‘agree’, ‘sign up’, and similar buttons, reviews the terms for you, and lets you proceed or cancel.
I used Opus 4.6 to build it. The extension uses Gemini 2.5 flash lite to analyze the terms
@appfactory@rauchg Why did you switch from Opus 4.7 to gpt-5.5? Are you experiencing any benefits?
I haven’t really tried Codex and I’m wondering if the switch is worth it
i present to u the most important NFT. RT for a chance to win one of ten exclusive #McRibNFT
no purch. nec. 50 U.S./DC, 18+ only. winners need crypto wallet to receive NFT. rules: https://t.co/2QRhsPlpur