here's how i shipped 2,500 PRs last month to production
this was originally supposed to be for Cursor Compile in London. i couldn't make it since i was livestreaming for Grok @Bot Galaxy so i'm making it available for free here on X! watch it on 2x speed, i talk slowly
An idea from Atomic Habits:
When you fall in love with the process rather than the product, you don’t have to wait to give yourself permission to be happy. You can be satisfied anytime your system is running.
mattpocock/skills v1.3 is out!
- /pr (new) writes easy-to-read PR bodies, showing hard evidence that the changes work and assessing merge risk
- /implement-spec (new) takes in a spec and tickets, and implements them with subagents
- CONTEXT.md renamed to GLOSSARY.md
- /retro (new) reviews recent coding agent sessions and suggests repo improvements
/retro, especially, feels like a huge upgrade. Enjoy!
You don't have to commit your entire life to optimism, just try it on for a week. Start thinking "what if it all worked out?" and "what a time to be alive!". Then compare your state of mind to the week before filled with pessimism and doomer nonsense.
Prompt of the day:
/writing-for-agents my AGENTS.md is a hot mess, propose a series of restructurings that:
- Remove no-ops
- Use progressive disclosure
- Move instructions to CODING_STANDARDS.md
Apply your work over three subagents, each more radical than the last. Create a single PR.
AGENTS.md is the most common source of token inefficiency. This prompt kills it dead.
We turned Claude Code, Codex, Hermes, Pi, @opencode and other coding harnesses into RL environments. No changes to the harnesses, no changes to the training code. Any open model, any task set, fully open source my friends!
Same model, same weights: 62% under Mini-SWE-Agent, 33% under Claude Code. But training inside a real harness normally means reimplementing it as an environment, so most models get trained in a scaffold nobody actually ships.
The fix is a proxy, not a rewrite. The harness thinks it's talking to a model API. It's actually talking to a capture proxy that speaks the 4 formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini), forwards to @vllm_project, records the exact token IDs and logprobs vLLM sampled, and hands TRL sequences it can train on. The harness becomes the environment. 10 harnesses run through it today, none modified.
And because you control the reward, you can shape behavior the harness never asked for. We added a small bonus for solving a task in fewer tool calls: on tasks it already solved, the model now uses 31% fewer calls, in every harness, and about half under Codex.
Tested on LFM2.5-2.6B from @liquidai:
→ Train in one harness: better mostly in that harness (OpenCode 34% → 58%).
→ Train in 4 at once: better in all 4 (42% → 54%).
→ SFT on 3,189 rollouts from Qwen3.8-27B instead: plateaus at 47.5%, below both RL runs.
Everything is open and reproducible: the capture proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, the training code and all 7 trained models. Bigger models and bigger runs next.
Full guide: https://t.co/sKZURuOcza
Software engineers are supposed to do exactly what we’ve always done. Ensure that the composition and structure of the system is robust and maintainable and that the behavior of the system properly serves the needs of our users.
This requires engineering discipline. It need not require detailed reading of the code; just as using compilers does not require detailed reading of the binary.
everything i know about managing agents i learned from the amazing programmers and computer scientists that came before
constraints in your codebase are freeing for both humans and agents! this has always been true, but if you were a small dev team you likely never had to think that deeply because you could trust your fellow programmers to write code and review well
but large companies have had to solve this problem since even before agents. because before you had agent slop, you had human slop. lots of engineers wrote bad code out of necessity (eg backend engineer forced to write frontend or reverse). the solution to this was constraints: lint rules, smarter compilers and diagnostics, high quality tests, investments into observability, and so on
the arrival of agents just means that big company problems are now everyone's problems. the good news is that none of this is really that novel: it's just good engineering
https://t.co/ykgu52qBqB
If it works,
if it's performant,
if it has no bugs,
if you understand how it works,
if the agent can modify it a month from now, to your liking,
is it slop?
people who laugh and say karpathy's latest tips are outdated - i suspect the vast majority of them have not even tried the tips yet
i just tested the ASD-STE100 wording rule and it's surprisingly good at helping increase clarity of model response, even within html artifacts. but the trick is that the full ruleset is a bit too strict and you need to pick a subset
one prompt you can run super easily:
"randomly sample 10 session transcripts where i worked with you interactively within the past week. apply ASD-STE100 rules to assistant responses and analyze which rules would have increased clarity, reduced confusion and improved the conversations, then document those rules in my user level AGENTS.md"
you might be impressed!
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
People seem shocked that the author of Clean Code doesn't read code anymore. But that should not be surprising.
Why did I write Clean Code? What is the goal of keeping code clean? To get it out of the way of the real job -- thinking.