Vendors who charge for tokens don’t have a pressing incentive to make their harnesses more efficient. As BYOK models and product evals get more popular, buyers will evaluate options based on performance - how much money do you transparently save them for the same outcomes?
This is a competitive advantage for software startups over frontier labs but not if they continue with their current pricing models, charging for inference rather than harness performance.
Having a lot of inference revenue might well become the innovator’s dilemma for AI native software companies - they can’t move to pricing models their enterprise customers will favor without destroying their existing revenue streams.
@PeterJ_Walker@Kellblog Nasdaq 100 grew over 6X in the last decade - giving away 20% carry for anything below that is a huge failure for LPs.. only your top 1% is good sir
@danielraffel@gokulr A juice cleanse for agents ;-) .. in general pre-loading too many things is a major cause of slow down for all agent harnesses - MCP, tools, skills - most of them should be invoked when required even if you do use them semi-regularly
Highly recommend that, was a very positive experience for me especially because my original files were created pre-memory. Minimal Claude.md, thin skills loaded only when needed, thin prompt with good artifacts and q&a for discovery is the way to go. I also kick off with Sonnet as fast default until the project is clear and then use better models for planning and orchestration - none of which you have to wait on.
So thrilled to try this out and couldn’t agree more - we need to expand eval rubrics from simple outcomes to include processes, especially for work where outcomes are not directly verifiable.
To paraphrase Bill Walsh, if your agents are following the right processes, the score takes care of itself 🙌🏼
Out of the box, long-horizon agents struggle to accurately perform end to end work in the real economy (outside of coding) because those tasks are not easily verifiable, the data is hard to scale, and going from inputs to real outcomes can actually take many days.
Even if you had a reliable way to verify outcomes at scale (and weren’t bothered by the multi-hour iteration loops), the sheer volume of decisions by the agent that occur in a multi-hour job makes it hard to know whether performing well will generalize to production.
Over the last two years at @trybasis, we've been solving this problem by supervising the process our agents take to get to outcomes, rather than just looking at whether the outcome itself is correct.
We think this is the key to building production agents at scale.
It's what has allowed us to run agents in production that operate for hours, sometimes days, and reliably perform tasks like entire complex tax returns end to end.
Today, alongside @braintrust, we're open sourcing a standard for defining, evaluating, and eventually rewarding agent behaviors.
Thread below with all the details on how we’re scaling behaviors to close the loop for long-horizon agents.
To understand and empathize with how workers in many or most fields outside software experience advances in AI capabilities, I propose a little thought experiment. https://t.co/QZG24aRiBe
@letsleverup Congratulations Aakash! Have often cited your thesis on revenue-alignment as a great way to approach vertical AI to my founder friends and clients. So excited for you working with Bret 🙌
Hardest IR of my career: one narrow objective, endless parallel paths, machine speed. One takeaway, we fought back with open models, in the open. AI security won’t be solved by one company in secret. Open source puts these tools in every defender’s hands
New USVC Investment:
We just bought a position in Anduril at the Series H price
America turns 250 this week and Anduril is our bet on America's ability to make things again
Here's why we invested in one of the most important private companies in America today:
@samhogan Super cool concept!!! Your approach to evals is worth calling out upfront though - most teams are not already happy with their current agent behavior (or shouldn’t be) so having some human calibration would make the results way more trustworthy
93% of technical execs we surveyed are alarmed about vibe-coded apps in their orgs.
Today we’re shipping the biggest update in @retool's history. One governed runtime for all of it.
There are some substackers/podcasters I liked following whose content has become claudeslop .. the substance and insight might be there but the writing is just so much worse than their previous bests and it makes me sad.
I feel an odd sense of loss and dismay as someone who loves reading well written thoughts.
@spenserskates hello Wave! Love this direction for Amplitude - honestly product data is the only way to push back on the deluge of slop code and slop features inundating all of us right now. It's great agentic coding is "solved" but we really really need agentic "product"