I just discovered one of those tiny terminal shortcuts that makes a surprisingly big difference.
Your server is running in one pane.
Tests in another.
Your coding agent in the third.
Then a long error gets squeezed into a tiny column.
In Ghostty on my Mac:
Cmd + Shift + Enter
The focused pane instantly fills the tab.
Press it again → your original layout comes back.
Nothing closes.
No dragging dividers.
No rebuilding your layout.
Useful for reading test failures, checking server logs, or reviewing an agent's output.
Here's what it looks like 👇
AI releases, graded as coffee ☕
GPT-6.1 Sol
Price: 1/5 of Astra
Taste: "near-Astra" on agentic coding
Verdict: flat white taste at office-machine prices.
But taste before you switch. "Near" is doing a lot of work. Run your evals.
GPT-6.1 Sol: $2/$10 per M tokens, 1/5 of Astra's price, "near-Astra" on agentic coding.
The real story isn't capability. It's that agent loops too expensive to run on every task just became viable.
But re-run your evals first. "Near" is doing a lot of work
Hot take: most "AI agents" in production should be 90% boring code and 1 LLM call.
Agents earn their place only when the path is genuinely unknown upfront.
If you can draw the flowchart, write the code. Use the LLM for the one box you can't define.
Python + SQL developers: how are you using AI to test your applications after every code change?
I’m looking for a setup that checks API behavior, database migrations, data integrity, permissions and end-to-end flows—and flags regressions before deployment.
Does AI help you identify missing tests, write them, or investigate failures?
What runs automatically in your workflow, and what still needs human review? Would love concrete tools and examples.
Your AI agent doesn’t need to reinvent the checklist.
Imagine asking it to check whether a release is ready.
Some steps are already known:
Run the tests. Collect failures. Pause for approval.
The part that needs judgment is understanding what failed and why.
GitHub’s new Copilot dynamic workflows let you put that process in code and bring in agents where analysis is needed.
For example:
• Code runs the test suite.
• An agent explains the failures.
• The workflow pauses for a human to review.
You can reuse the same process on the next release. The agent’s answers can still vary, but the required steps don’t have to depend on what it remembers to do.
A useful design question: which decisions actually need AI?
Available in public preview:
https://t.co/bUEv5dOQLv
Your AI fixed the test. Did it fix the bug?
Imagine asking a coding agent to fix duplicate payments.
Before:
Retry the request → two charges.
Test expects one → fails.
After:
Retry the request → two charges.
Test now expects two → passes.
Everything is green. The bug is still there.
When reviewing an AI-generated fix, check the test diff:
• Did an assertion disappear?
• Did the expected value change?
• Did a mock bypass the behavior being tested?
Changing tests can be correct. But the reason should come from the requirement.
“Tests pass” is useful evidence.
“What do these tests still prove?” is the review question.
Claude Code can now help you rebuild Claude Code.
Anthropic just introduced “mods”: small TypeScript functions that change how it behaves and looks.
Imagine your agent reads a log containing an API key.
A mod could intercept the tool output and redact the key before it reaches the model. That’s one practical example.
Mods can also:
• Block, rewrite or retry tool calls
• Add buttons and custom panels
• Replace built-in features like /diff
They work by running before, after or around events inside Claude Code. They ship as plugins for the terminal and desktop app.
A useful starting point: ask Claude to build a mod that displays your build status beside the conversation.
One catch: mods aren’t sandboxed. They have Claude Code’s access to your machine, so review what you install.
Your coding workflow can become a feature you build.
https://t.co/QIELGbq2WL
@Gaus450 Whichever one needs the weaker harness to do the same job.
The interesting benchmark isn’t model vs model anymore.
It’s:
model + context + tools + retries + verification + cost → reliable outcome.
Claude Code: undo the bad code. Keep the conversation.
A refactor fails. You explain why.
Now you want the original code back without losing that discussion.
Run /rewind → select the checkpoint → “Restore code”.
The tracked edits roll back. Your conversation stays.
Then ask:
“Try a simpler approach using what we just discussed.”
Useful for failed experiments when the diagnosis is worth keeping.
Keep using Git: rewind doesn’t undo shell-command changes or most subagent edits.
@dhh The scarce resource isn’t speed. It’s generating hypotheses worth testing and killing the rest before they compound. Acceleration only works when evaluation is cheaper than debate.
Distribution. Technical skill builds the thing. Sales skill gets the first customers. Distribution is what turns both into a company instead of a project.
You can hire engineers. You can hire closers. You cannot easily hire the founder’s ability to put the product in front of the right people, repeatedly, at increasing scale. Most great products that failed, didn’t fail on quality. They failed on reach.
@AravSrinivas What you wanted as a kid wasn’t a video. It was someone who could hold the whole system in their head and walk you through it without getting tired. That’s the actual product.
The interesting part is triage: did the API fail, did the caller misuse it, or is a capability actually missing?
For bugs, have the agent reproduce the failure before proposing a fix. For features, define acceptance criteria first.
That gives the human reviewer something concrete to approve.
@NickpxJ I would hire a developer who uses it well.
When an AI-generated change passes every test but charges a customer twice on retry, I want someone who can spot the missing requirement—not just write more code.
The best Claude Code upgrade isn't a better prompt.
It's giving Claude a way to know when it's wrong.
Give it:
→ tests for backend code
→ a browser for frontend
→ linters/type checks for code
→ validation scripts for data
Now the loop changes:
write → verify → fail → fix → verify
instead of:
write → “looks good”
Anthropic calls verification its #1 Claude Code tip.
A smarter model helps.
A model that can check its own work helps more.