I let AI write most of my code now.
I'm Eric — I build with Claude Code & Codex every day.
This account is my public notebook:
→ real workflows that actually work
→ honest tool comparisons
→ failures (plenty of those)
No hype, just practice. Follow along if you're building with AI too. 🛠️
New study from Singapore Management, Shanghai Jiao Tong and ByteDance takes a torch to one of the favorite agentic-coding selling points: an agent writing its own tests barely helps it fix bugs. On 500 real GitHub bugs, the ones agents fixed and the ones they didn't showed nearly identical test-writing habits. And when researchers prompted agents to write more tests, no model's outcome shifted in a statistically significant way.
@gdb Once the UI can render anything, a text-only answer starts feeling lazy. The next skill agents need isn't better code, it's better taste: knowing when to draw the interface instead of describing it.
@bcherny At 10x cheaper with 100k context, Haiku earns the whole fast loop: planning, review passes, small refactors. I only want the expensive model spending tokens on the parts that actually need it.
It makes sense once you see it. The same model that misread the requirement also writes the test that fails to catch the misread. A print statement cannot block a bad fix. An assertion can. The tests that matter are the ones the agent did not write. Paper is "Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents", on arXiv: https://t.co/js8VjqGza1
My favorite detail: GPT-5.2 fixed 71.8% of the bugs while writing new tests in only 0.6% of its tasks. And when agents did write tests, they wrote print statements, not assertions. For Opus 4.5, prints were over 82% of the feedback. They were not verifying anything. They were peeking at values and eyeballing them.