@evandrocabf@spenserskates Great question! Our OOTB enrichments are based on the premise that when you start an agent interaction with a goal in mind: competing a task, creating content, random stream of consciousness, etc… Would love to hear your thoughts on this
Today we're launching Agent Analytics
Every team shipping an AI agent has the same blind spot. Offline evals pass, you ship, and then you have no idea what's happening in production. AI fails silently. Users ask a question and get different answers. They all look 'engaged' in a classic dashboard. You don't know who got a great response and who got a terrible one.
Agent Analytics solves it:
- Every session scored out of the box on task completion, response quality, friction, safety, and negative feedback
- Topic clustering across thousands of conversations, so you know if a failure hits 1 user or 10,000
- Eval agents that watch for regressions, and if you want will file a Linear ticket or the pull request themselves
- Agent quality sits next to product data, so 'payment scheduling fails 31%' becomes 'which renewals did that cost us?'
The Economist got their agent to a 96.9% task success rate and cut weekly failures 84%.
Included on every plan. Free tier included. https://t.co/K5vqrlMLvd
@theo We’re working on a lot of this actually. You can send us your agent traces and get insights directly from Claude via our MCP. We’re starting to pick up a lot of traction, give us a try and I can share your feedback with my team.
https://t.co/IKmtAm6yhU
That's why we built Agent Analytics, available today on every Amplitude plan. It scores every session automatically (task completion, friction, user safety) and clusters them by topic, so you can see exactly which failures are costing you a renewal. Stop shipping on vibes only.
A passing eval score tells you nothing about whether your agent is helping your business. No company builds an agent just for the sake of it, they want conversion, retention, monetization. Offline evals work fine against a ground truth dataset, but production is a different game.
Software velocity compounds. Every PR is an experiment, the faster you run them, the faster you learn.
Six months of optimizing our engineering system:
- 3x PRs
- Cycle time down 7x
- Bug reports down 55%
- 5% of PRs from non-engineers
We turned the whole story into a game so you can speedrun this too 🕹️
Enjoy the run!
Your work tools in Claude are now available on mobile.
Explore Figma designs, create Canva slides, check Amplitude dashboards, all from your phone.
Give it a try: https://t.co/hwPB3zlk0w
How do you ship AI products reliably? 🚢
Join Amplitude’s @JacobNewma22350 and @ilankirM on March 20 for a behind-the-scenes look at building agents, designing for trust, and using evals as a living product spec.
Save your spot here: https://t.co/0isbnVyiA0
excited to share: @cursor_ai 🤝 @Amplitude_HQ
with the amplitude plugin, i can pull product context directly into cursor to analyze dashboards and synthesize customer feedback, then push requests to cursor agents to draft PRs.
when those fixes ship, we track usage and feedback in amplitude, which feeds right back into cursor.
call it the first self-improving product loop 🔁
Who else spends a good portion of their day in tools like @Claudeai and @NotionHQ?
With MCP, now you can access @Amplitude_HQ ’s behavioral context as part of that workflow. Ask questions, investigate issues, and act on what you find without context-switching.
#AmplitudeAI
Al has collapsed the time it takes to ship a feature. Now it's all about shipping the right thing. It's just as hard as it ever was.
As coding gets automated, agentic analytics becomes the bottleneck. Teams that win will build within context rich environments that capture user behavior and customer feedback, making that data accessible to both builders and agents.
Today @Amplitude_HQ is launching our Al Analytics Platform with Agents and MCP. Autonomous analytics that helps builders and coding agents ship products that their customers actually want.
Just as our partner @claudeai has reinvented software, Amplitude is reinventing analytics for the Al age.
We have achieved 76% accuracy for complex production grade queries and agentic usage has grown 10x in 3 months.
Don't just build fast, build what's right.
Try it for free, today.