One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
In the world of AI agents, I’m realizing that understanding why a metric is going up unexpectedly is just as important as understanding why it’s going down.
Example: today, our agent containment rate increased by ~10 percentage points without any changes on our end.
The easy thing would have been to chalk it up as a win. But after digging in, we realized the increase was actually caused by something that negatively impacted the end user experience.
It’s a good reminder that with AI agents, higher containment is not always better. Metrics can move in the “right” direction for the wrong reason.
@vasuman Sounds like a dream. Uninterrupted focus on getting the AI agent to actually work. Takes a lot of chipping away. Existing folks who can probably do it at their companies also have a day job
I don’t think he’s saying chat itself is a demo, more so that chat for booking travel isn’t solved yet. Sounds like he has an insight that a large % of this users are booking Airbnb for group trips (which resonates) and maybe he’s building for that main use case these days. I haven’t used Airbnb in a while but this was my old exp:
1. Find Airbnb
2. Send links to friends
3. Discuss pros/cons with friends
4. Find more options
5. Repeat steps 2 through 4
6. Decide and book
Would be a crazy concept if there was a group chat for AI
Totally agreed. And when I say “monitoring,” that’s what I mean. I’m staying in the loop, reviewing evals/ edge cases, but I also find myself dogfooding and spot checking trace logs, reviewing conversations, etc to come up with more evals and ways to fine tune or handle other scenarios. Hours can go by pretty quickly doing this :) it’s an interesting realization - that maintenance of these agents quickly becomes the most time consuming part. Vs old days where “building” a feature was the most time consuming
Your key point is “what the outcomes were.” It’s so easy to measure the impact of layoffs on balance sheet from a cost perspective. But I’ve always wondered how one can measure the negative impact - lost revenue, tribal knowledge, lost opportunity on roadmaps, etc. but who knows maybe these companies truly were bloated to begin with