"In a time of drastic change it is the learners who inherit the future. The learned usually find themselves equipped to live in a world that no longer exists." — Eric Hoffer, 1973
Evals have been coming up more and more in my conversations with podcast guests and PM friends.
Nearly half of the 25 awesome PM job openings I shared last week ask for experience writing evals. And leading companies keep sharing what investing in evals bought them:
— @tryramp took its automatic receipt collection from 35% to 83% accuracy.
— @Shopify shipped an AI workflow builder that's 2.2x faster and 68% cheaper than the frontier-model setup it replaced.
— @harvey__ai rebuilt its AI contract reviewer, nearly doubling its internal quality score.
— @cursor_ai tuned its Auto Balance routing, with much higher user satisfaction at 41% lower cost.
So I asked the 🐐s of evals, @HamelHusain and @sh_reya, to write an advanced sequel to their very popular "Building eval systems that improve your AI product." Drawing on their work with 50+ AI companies, they share the key step most teams skip, what you should (and shouldn't) automate, and a free plugin that lets a coding agent do most of the heavy lifting.
Read it here: https://t.co/fxOI7PMgWF
The problem with all of these AI personal assistant things is that people don’t actually want to do stuff.
You fucking get that right? Do you understand that?
Just like AI tools in the workforce, AI assistants don’t actually abstract decisions for you, they let you accelerate decision making.
So instead of making one decision a day and then lamenting you can’t do more, AI assistants are going to let you make ten decisions a day and act on all of them.
A miracle?
Nah dog, a fucking nightmare. Nobody except type A strivers who need Twitter fodder actually want that.
Let me tell you a story.
When I got back from college I was inspired. If I could take on an elite university with grit and determination, so could everyone I know.
I was all razzed up to get people I know to achieve their potential. What I learned after many years of trying is that absolutely nobody wants to realize their potential.
People want to hang out with friends and complain. That’s how the world has worked for ten thousand years. This weather fucks. Can you believe the cows walked off.
People don’t want enablement. People don’t want you to take all of their excuses.
TV is mindless passive activity. AI assistants are not. They’re telling people that they can achieve more. And people absolutely do not want that.
Let me tell you another story.
Once upon a time I lead a big process overhaul at work and it was great. Two years later I hired a contractor to do something similar. But I personally didn’t want to have to do a bunch of work.
Within a week I knew I fucked in. This contractor kept coming to me with decisions I needed to make. Dude I wanted to not have to think about this.
Again, a helper when you don’t actually want help is a disaster.
None of these AI assistants will take off because people don’t actually want assistants.
I am thrilled to announce that @beaconholdings has acquired @haizelabs, with me joining as VP of AI Research.
We started Haize in 2024 to enable anyone to build reliable and safe AI. Through our red-teaming and safeguards work with the frontier labs; our observability, guardrail, and evaluation platform serving the world’s largest enterprises; and our pro bono SMB work, any customer could build AI they trusted with Haize.
Joining Beacon lets us deliver the same trustworthy AI to those who need it most, and those most overlooked by Silicon Valley: the essential Main Street businesses the real world depends upon.
Transitioning these essential businesses through the AI revolution is one of the most consequential responsible AI problems of our time. We couldn’t be more honored to tackle it with Beacon.
Thank you to our customers, investors, and team for the journey of a lifetime. And thank you to Nilam, Goutham, Mark, and the Beacon team for the trust and opportunity.
It’s time to get to work.
It's isn't just to people with ADHD that this happens. Some tasks inherently take a certain amount of time. An appointment that breaks the afternoon into two blocks of time could leave you with no block long enough to complete a task.
some underrated points + important second order effects of the wildly successful Jev launch
1. Data is the bottleneck! TypeSafe calls themselves a “data research lab”.
To build a general purpose classifier or actor of any kind, you need to painstakingly curate and create tons of high-quality data (Traces, Examples, Evals/Environments/Worlds, Simulations).
This helps the model learn these distributions so it can do good work for users. This is especially true for building domain specific, vertical agents.
Look at the data. Every team looking to build better agents will need help + tooling to help them with this. If teams can spend more time building good data, then they will build better models & agents.
2. The explosion of ultra-cheap, always-on Monitoring + Data Mining of Agent Traces at scale
To build better agents, we need to understand them at scale. We also saw the large need to real-time monitor and flag bad behavior during the OpenAI-HuggingFace Incident. Models like Jev help us do this because they’re very fast and cheap for classification.
There’s a simple pipeline here:
Gather Traces —> Label Them aggressively —> Make Evals/Environments from them (optional) -> Fine-Tune Cheaper Model.
Part of data hygiene is understanding data at scale. We sort of had the tooling to do this before by fine-tuning small models or sending smart agents to read lots of traces.
We’ll still use agents to mine traces because Jev has limitations like context window length, but this is a great tool for:
- humans to define dimensions up front on what they care
- using Jev to triage data across important dimensions for further review.
3. We are never escaping Jevon’s paradox
Models like Jev just help us do more. Classify more traces, monitor more actions in real-time. This is all net-new addition of compute and will happen at massive-scale over every piece of data agents create.
This is all very exciting because though we’ll use more compute, we’ll understand much more to build better systems.
Played around with Jev before I went to bed and I'm really impressed. It also fits so perfectly well for so many applications where traditional LLMs so far were just not viable for either speed or cost reasons. I bet we will see some fast followers. https://t.co/FaZvpo1JI7
@GergelyOrosz I think death of middle management may need to be evaluated in terms of switching from managing people to managing agents. I think going forward everyone is a manager actually. But some folks will turn themselves into vps and ctos of agents.
I can’t stress enough how little an idea matters compared to the agency of the people executing the idea.
I have had the privilege of knowing and sometimes even working with some of the most successful people (by various metrics).
The difference between mediocre and excellent work and outcomes is predominantly one of agency.
In practice this means: they dont wait for things to happen to them they go out and make things happen for them.
They don’t wait for someone else to do something, for someone to teach them, for someone to give them the path, etc. They just go out and find a way to do it.
I think the single biggest superpower these people have is the realization/belief that the world around them is completely mutable. Most everything that happens is because a person made it happen.
I used to tell people to look around the room you’re sitting in. Look at everything. Every noun. It almost all exists because a person willed it into existence. Nothing is stopping you from doing the same.
I see people online all the time dismissing someone else’s success because “I had that idea first” or whatever. I mean… yeah? If so then the difference is… you. So a bit of a self own whenever I hear that.
Number one tip: act with agency.
I’m one of at least 28 former Washington Post journalists who have joined the @washingtonsun, which launched this morning. Paul Kane, Jeff Stein, Tom Sietsema and many other names will be familiar to Post readers —and we’ve got the Capital Weather Gang. Join us! https://t.co/qK806amX3p