Yesterday I shipped Pitwall, an open-source Claude Code plugin that routes coding tasks to different model-harness combos and verifies the result.
Today: run your own bake-off to find the right model and harness for your work, with dollars per task as the north star.
A/B the same request across two combinations, or sweep every one you've configured. It reports cost, time, and whether the result passed your gate.
Leaderboards rank models in the abstract. This ranks them on your repo, your task, your definition of done.
My run surprised me. Harness choice alone moved cost by close to an order of magnitude on the same model. Every orchestrated lane cost more than doing the task in one plain session. The pricier combinations bought no better output.
Yours will differ. That's the point.
Open source: https://t.co/bIc2aoYFmU
After months of playing with model and harness combinations, one shape kept winning.
A frontier model for planning. Model-harness combos, efficient in dollars per task, for execution. The reason is subtle: the frontier model does the loop engineering for you.
So I built Pitwall, an open source Claude Code plugin that puts this into practice. Inspired by Omnigent. It's also the first thing I've shipped seriosuly in well over a decade. Building is fun. https://t.co/Pq0NiXGqlj
This is one of those unintuitive things. Agents that cook longer are often worse. Genie just gets to the results faster. Ontology will be key to getting these agents the context they need to get the answers right quickly.
After months of playing with model and harness combinations, one shape kept winning.
A frontier model for planning. Model-harness combos, efficient in dollars per task, for execution. The reason is subtle: the frontier model does the loop engineering for you.
So I built Pitwall, an open source Claude Code plugin that puts this into practice. Inspired by Omnigent. It's also the first thing I've shipped seriosuly in well over a decade. Building is fun. https://t.co/Pq0NiXGqlj
We're raising funding at $188 billion valuation to double down on our AI strategy focused on three priorities:
1️⃣ Unity AI Gateway - our multi-AI governance solution that helps control costs.
2️⃣ Genie - our AI coworkers that actually understand your business data.
3️⃣ Lakebase - our serverless Postgres database specifically for AI agents.
https://t.co/Bnc89qxjLG
A paper diary that answers back.
You handwrite a question. The ink fades. An answer writes itself onto the page - drawn from your own notes, not the internet.
It's an AI grounded in my own notes, living inside a reMarkable tablet.
A paper diary that answers back.
You handwrite a question. The ink fades. An answer writes itself onto the page - drawn from your own notes, not the internet.
It's an AI grounded in my own notes, living inside a reMarkable tablet.
Why do people who agree with you completely still refuse to change?
You are talking to the rider. The elephant is what walks, and it has not felt a thing.
How to reach it:
https://t.co/BJwjuJOgjS
1/ Information almost never changes behavior.
The people who fail to change are usually the ones who already agree with you.
I spent a while working out why, and what actually moves a person. A thread on the engine of change:
14/ None of this is manipulation. It is the difference between talking at the rider and reaching the elephant.
The arguments are usually already there. The only real question is where the thirst is, and whether anyone has made it felt.