Part of the LLM-assisted work here is that Fable+Sol put together a test suite of over 400 cases that are bare shell scripts plus a harness that runs Kitty and Ghostty GUIs and captures their full pty stream AND screenshots (multiple, for animations).
For the pty stream we assert byte equality. So Ghostty's success and error messages and field order of responses (its k=v) directly matches Kitty byte for byte.
For screenshots, we can't do pixel-identical comparisons because the way Kitty and Ghostty calculate grids and do alpha blending doesn't match. But, LLMs are pretty good now at "do they look the same" PLUS I went through all the screenshots myself in the end.
Super helpful AI assist. It allowed me to focus a lot of my brain energy and time on reading the spec, reading/writing the implementation, reviewing a lot of code, considering the right shape of things, the performance implications, etc. while I had a couple very good interns in the background doing work like this.
I plan on open sourcing this validation set and harness, with the full disclaimer that it is 100% AI written. But, I think its a perfect example of something that SHOULD be.
The way I had Fable+Sol work together here:
1. Sol put together the harness.
2. Fable + Sol (two separate agents) in parallel would write test cases and output them on disk in their own folders. These two are just in a ralph loop.
3. Sol + Fable (reversed) with an adversarial prompt would judge the others work by picking up the changes on disk that step 2 wrote. They would determine how accurate/worthwhile it is to keep. They'd put it in another folder.
4. Sol finally woke up for changes to this final folder and would determine if its a dup or not and then add it to the final repo.
Then I'd pick up the bug reports, validate them myself, and either fix them myself or kick off new agents manually, just Codex app or Claude app.
Finally, re-ran both agents once against the full test suite to verify what I saw myself: everything passed, all images look the same.
@jessethanley@nateberkopec We’ve been automatically approving changes to tests and fully merging dependabot updates for years. Now we have agents approving safe PRs: https://t.co/kxahI4PoJv
Once you reframe it around safety, you can get a pretty high number of PRs fully approved.
@GergelyOrosz I went through this journey myself and I tend to split them into three categories:
Product loops - working through a set of tickets or todos
Reactive loops - triggered by alerts, typical oncall stuff
Proactive loops - maintaining quality or refactorings
I have been hyping @brian_scanlan and @intercom to every eng leader I've spoken to since getting a behind the scenes look at how they 2x'd their productivity in their R&D org.
In this week's episode of How I AI, Brian pulls up dashboards and tools and shows us how it's done:
- explicit targets + reporting on leading indicators of AI adoption
- how they measured how quality increased with more AI usage
- centralized skills repo for the entire company
- telemetry on claude code + evals on internal success
- building for an agent-first onboarding flow
Sorry to be that bro, but truly: STOP what you are doing and watch this right now if you are looking to inflect your eng + product team with AI.
As always, a huge thanks to our sponsors:
🧠Celigo - Intelligent automation built for AI: https://t.co/wXLyXeR53y
@cursor_ai - The best (and my fave!) way to code with AI: https://t.co/GBqDSfjHdu
On YT now 👉 https://t.co/YLItD2lKFc
9 months ago we publicly committed to 2x the productivity of our R&D org at @intercom.
It was scary. It wasn't always clear we'd pull it off.
We hit it with 3 months to spare. In fact, looking back 16 months - we've 3x'd.
Here's what actually happened (with receipts): 🧵
Orange Brompton C Line 6 Speed stolen today on Chancery Lane, London.
Low handlebars (flat bars) bike Is orange with black wheels and fittings.
I changed the bar grips to the foam type and it is fitted with Brompton Be Seen lights front and back. Serial 2409110488
@StolenRide My Ribble CGR AL Sport was stolen from my buildings bike storage today in Stamford brook.
It's a medium size frame and has Deda Zero 1 Stem, Deda Zero 1 RHM Bars, Fabric Line Sport Saddle and Shimano EH500 Pedals that are not stock with this model
PICARD SHOULD DROP ALL OF THESE FAKE POLITICAL INDICTMENTS AGAINST ME, BOTH CRIMINAL & CIVIL. EVERY CASE I AM FIGHTING IS THE WORK OF STARFLEET & BAJOR. NO SUCH THING HAS EVER HAPPENED IN THE ALPHA QUADRANT BEFORE. JUMJA STICK REPUBLIC??? ELECTION INTERFERENCE!!!
This isn't just good leaflet copy, this is why politics matters and different parties aren't all "just the same".
Here's a collection of similarly devastating graphs
🧵 (1/11)
@moultano This isn’t my experience in Asia, lots of “expats” in HK and Singapore quite happy to have live in maid(s). They usually have some BS about treating their maid better than local but they are still using cheap labour and shitty conditions.
not saying the last year has been chaotic but in this time Grant Shapps has been home secretary, transport secretary, business secretary, energy secretary and now defence secretary