We are releasing the world's first case study of self improving products at Amplitude. A lot of companies have talked about it but none have shown any examples. We have the holy grail of this space and can show it!
We've been running our self improving product engine, Wave, on top of Amplitude for the last few months. It's made lots of changes to our website and product based on analytics, session replay, and many other data sources. These have resulted in MASSIVE improvements in metrics on the order of 50% to 3x, not incremental ones.
These include:
- Personalized suggestion cards on a homepage led to a +160% increase in usage
- Improved doc search accuracy from 92% to 97% by setting a 3 character minimum
- Reducing setup checklist from five next steps to one guided first step, projected to improve onboarding completion from ~6% to ~9%.
- Adding a CTA to each blog post leading to 45 extra signups a week
We're slowly rolling Wave out to a few customers and will be sharing examples from them soon. Check it out here and sign up if you're interested in entering the closed beta:
(1/6) i'm the pm on Wave at @Amplitude, our agent for self-improving products. for the past few months we've had it running on our own product. it finds problems, ships fixes, and measures whether they worked. here are some highlights:
When it comes to the SF Fentanyl crisis, the problem isn’t the government, it is us.
We’ve told our government to arrest drug dealers and house/treat drug users.
In the past 3 years we have over 1700 dealer arrests, almost 5000 people given housing, and 25% increase in treatment (and we have excess supply of treatment).
Unfortunately that strategy hasn’t worked. And it won’t work.
The drugs are too cheap to produce and transport and the dealer labor supply is infinite.
Giving a user housing doesn’t magically cure their addiction and scaling drug treatment in places where drugs are easily available and cheap is highly ineffective.
This is a complex problem and we need a more sophisticated solution.
Everyone points to Zurich as the best case study. But advocates put too much emphasis is on safe consumption sites and almost none on the role of the police.
We are going to have to better educate ourselves about the problem if we are going to expect better solutions from our government.
AI makes this information readily accessible.
The rate of progress in the labs is accelerating
The rate of adoption of that progress in society is not
It’s going to go both very fast on one end and it will feel slower than ever to those who understand AI because society has natural brakes
Fast takeoff with slow uptake
The system is set up to extract wealth from Millennials and Gen Z and give it to Boomers:
- NIMBY housing laws that make it impossible for first time homebuyers to afford a home.
- Student debt that can't be paid off or discharged, leading to permanent wage depression.
- State governments in unsustainable levels of debt to civil service pension funds while canceling those benefits for new hires. California pays an average of 102% of final earnings to retiring civil servants.
- Proposition 13 lets longtime homeowners pay significantly less property tax than new ones.
- Both Medicare and Social Security will not be fully funded starting in 2033.
- Politicians are terrified of calling this out because 9 out of 10 people who go to their town halls are Boomers. I talked to one US Congressman who said "coming out too hard against Boomers is an almost sure way to lose an election". They're happy to call out billionaires but terrified to call out Boomers.
Once you see it the first time, you see it everywhere.
One corollary: have fewer managers overall
We've grown flatter at Amplitude. We have fewer managers at Amplitude than we did 18 months ago, even though the company got bigger.
I pulled the stats from start of 2025 to today:
- Managers: 158 → 135
- Individual contributors: 567 → 656
- Managers have gone from 3.6 reports per manager to 5
When you talk to successful founders, they always wish they'd hired fewer top managers from outside. I don't think I've ever met a founder who wished he or she had promoted fewer people from within.
While the efficiency of models is improving quickly, it's not as simple as swapping out a cheaper model.
We tried Gemini 3 flash instead of Sonnet 4.6 on Global Agent and while the top line eval pass rate was similar for an 80% reduction in cost, latency was significantly higher and engagement was 10% worse.
Kimi K2.7 became competitive after targeted engineering fixes. It initially scored 64.0% versus Sonnet’s 72.7% on our eval suite. We identified failure cases including hallucinated capabilities, repetitive tool loops, calculation errors, chart mistakes, and output leakage. After adding tools and safeguards, Kimi scored 73.7% on the same cases. Once we ran it on a live experiment we got similar or better engagement for 52% less cost.
See how we did it with Amplitude on Amplitude:
Your AI bill can drop 80%.
Most teams get there by swapping in a cheaper model, and their users pay the difference.
We measured it on our own agent: half the cost per user, 88% slower responses, 10% fewer messages.
Here's how to cut the cost and keep the quality ↓
@vixsheikh We’ve been able to execute much more quickly is the main improvement. Each layer adds a 10x execution slowdown.
Main challenge is not as much redundancy