What if every escalation came with proof?
Launching Tier2: an AI support engineer that works your queue on its own. Read-only access to your infra over MCP. Simple tickets handled unattended. Engineers paged only for verified, reproduced problems.
2:16, real product.
engineers: what's the oldest unreproduced bug in your queue right now? the one nobody can make happen again, that everyone stopped believing but nobody will close. how long before a bug becomes folklore?
watched a repro land in a channel once with the commit author tagged by accident. the room went quiet. now the drafts name the change, never the person. 'introduced in 4.2' starts a fix. 'you broke it in march' starts a defense.
five customers hit the same bug and the queue treats it as five separate mysteries. reproduce one and the other four stop being mysteries. dedupe by evidence, not by reading tickets harder.
design decision: some bugs would repro faster if the agent could write to customer systems. it can't. read and probe only. a failure you caused yourself proves you can break things. it doesn't prove the bug was already there.
@TheArslanNasir no paste step. the draft links the run inline, approver reads the claim next to the actual probe output. a run stays good as long as the setup it ran against does - same version same state, replay it tomorrow. infra changes, we rerun and flag the old one
most support ai is sold on deflection. we took the opposite bet: the hardest escalations go to the agent first. deflecting an easy ticket saves two minutes. a bad escalation costs an engineer their afternoon.
@amiiinghasemi no - a run stays pinned to whatever it ran against: version, commit, state. when the infra changes we re-run against the new setup instead of rewriting the old run. stale runs get flagged on read, never deleted. old runs are evidence, expiring them would defeat the point
most support knowledge dies in slack threads. ours is executable: every claim the agent makes cites the probe run behind it, and the run replays. institutional memory you can re-run beats a wiki nobody reads.
reproducing the wrong bug is worse than no repro. now the fix chases a failure the customer never hit, and the real one is still out there. a repro only counts if it matches the report. close enough is a guess with extra steps.
every handoff resets the clock. new engineer opens the ticket, asks the questions the last person already answered, waits a day. a reproduced case is the only handoff that arrives intact: setup, run, log line, commit.
a bug never arrives whole. it's a slack thread, a cropped screenshot, half a stack trace in a loom. before you can reproduce anything, you have to assemble the bug.
security reviews never ask about the model. they ask three things: where does our code run, what can it reach, what can it change. isolated sandbox, nothing but its own deps, read and probe only. that's the whole call.
early mistake: we let the agent suggest fixes next to the repro. engineers spent the whole review arguing with the fix and skimming the evidence. now it reproduces, pins the commit, and stops. the fix is your job. the proof is ours.
most users never file the bug. they hit it, shrug, and quietly leave. the one who writes 'it broke, no idea why' is the only one who gave you a chance. vague tickets aren't noise. they're the last signal before churn.
watched an engineer review an agent draft. run, log line, commit, all attached. she changed one sentence before approving: the one that overpromised. proof can be automated. knowing what not to promise can't.
@TheArslanNasir rejects happen - usually a claim with no run behind it, the draft said something the probes didnt prove. and we never get to the write at all, read and probe only. approval stays about the evidence, not about trusting the agent anywhere near prod.
bug reports have a shelf life. logs rotate, pods restart, caches flush. the state that made it fail is evaporating while the ticket sits in triage. repro fast or the environment takes the answer with it.
design decision: the agent never says fixed. it says reproduced, or it says couldn't reproduce. fixed is a human call made after reading the run. an agent that declares victory is just generating adjectives.
engineers: what's the weirdest root cause you've ever pinned a bug to? the ones that sound completely made up are always the real ones. dns cache that only lies on tuesdays type stuff.
people ask what happens when our agent gets it wrong. worst case: a wasted sandbox run on a bug that was never there. it can't touch your systems, read and probe only. the failure mode is burned compute, not a broken customer.