This year we’ve all heard stories about blowing past AI budgets, allowing agents improper app access, and fighting against drifting output quality. The solution? Merge for Workforce.
We just cut AI spend by 75% across an entire workforce in one click.
Introducing: Merge for Workforce.
Without Merge, your employees are left choosing one of two options:
1. Spend 100x the cost they should be from using frontier models for everything
2. Deliver subpar work from overly restrictive model access
Instead with Merge, IT can connect any identity provider and set model routing policies by team in one click.
Merge pushes them to every machine through a desktop client, overriding the model configs across your team’s AI assistants and coding tools.
Now every task gets the right model for the job.
The result? 75x fewer tokens with faster and better output.
A lot of companies founded pre LLM explosion have struggled to find PMF by tacking on AI features that customers may not care about or want.
Check out what Merge has done to build meaningfully in the space - you'll learn a ton from @shensi!
https://t.co/dKlpuV3OQC
New @ThePeelPod with @shensi
Most founders get one founding story. Shensi thinks @merge_api has three.
Shensi and @GilFeig started Merge in 2020, watched their first customers die, then rebuilt themselves twice for the AI era.
Not many founders talk this honestly about tearing up their company. A lot of lessons for others trying to do the same.
We get into why integrations turned out to be so important in AI, the playbook behind their night-shift 5-11pm AI transformation, almost hiring a foreign spy, why founders with an EA are moving too slow, the Embarrassment Framework, vibe coding a dinner bot that 10x'd their customer events, and why 6% of Merge employees get married.
Full episode here + links below!
Timestamps:
0:00 Something broke every year since 2020
2:22 How Merge went all-in on AI on nights and weekends
5:40 New launches got faster
8:32 Building connective infrastructure for AI
10:51 The dinner where Merge started
13:11 Six months of research before a line of code
19:00 Almost hiring a foreign spy
20:48 Merge’s 6% marriage rate
22:50 Hiring enthusiastic, nice, smart people
28:02 Early stage founders don’t need an EA
31:27 Shortcuts are a mentality
33:01 The AI tool that 10x'd their customer dinners
38:33 Marketing became an engineering function
40:51 Starting with SMB and climbing the logo ladder
42:07 How the product went cross-category
44:16 Launching Agent Handler and a new pricing model
46:47 Every product should be multi-model
49:58 Why most MCP servers don’t work
53:31 Startups should go all-in on enterprise
57:00 The hardest things are most defensible
1:00:26 Say "psycho shit" to be memorable
1:05:47 The Embarrassment Framework
1:10:23 Frank Slootman and being okay with being disliked
1:12:48 Giving feedback got scarier at 100 people
1:14:01 What only the CEO can do
1:19:44 The best marketing is not doing what everyone else does
@shensi@ssh_sharma135@FactoryAI@merge_api Wait, this is giving me ideas to do a Merge ice cream x cafe pop up. Small batch flavors, custom affogatos, pour overs 👀 🤩 😋
@Suhail Check out @merge_api! We've even run evals showing how our intelligent routing can save nearly 2/3 in spend at virtually no change to output correctness 🤑
@zaingz@hazrmard Yes, online cases from users are tough since logging may be off due to PII or confidential information. Curious to hear what you or others can do to build around this constraint
@hazrmard Interesting angle on (1). For offline evals and regression tests, we have a fixed prompt set that's meant to be the golden dataset. More capable agents and models that actually outperform the rubric or find holes would be something to flag for rubric self improvement
If you've made it this far down, hopefully you get a sense for how nuanced evals work is. It's foundational to product quality in this AI era and the only way to know if your offering is good in production instead of some nice looking demo.
Much of my career in data has been understanding what people are actually trying to do. The same applies here: ground eval prompts in real use cases and failure modes, not hypotheticals. Customer feedback, support tickets, and dogfooding beat guesswork.