Coinbase halted trading for about eight hours on 7 May. Twenty hours to full recovery.
AWS lost a zone. Their matching engine was pinned to that one zone, and the managed Kafka they'd paid for to survive exactly this had a control-plane defect.
Managed isn't immune.
You can improve your p99 without fixing anything. Lower your timeout.
Requests that time out become errors and leave the latency distribution entirely. The number goes down. The experience doesn't change.
Every latency metric has a survivorship problem.
Monitoring now averages 17% of infrastructure spend. 37% of teams name cost as a major concern; 74% weigh it when choosing tools.
Grafana's survey, 1,255 responses.
A monitoring bill you have to justify quarterly is a monitoring bill you eventually cut.
"Production bugs are like shadows. They lurk until you shine a light with observability. That's where https://t.co/wUf1LhV2pD comes in."
"Debugging is the art of understanding chaos. Let https://t.co/wUf1LhV2pD help you tame it and keep your production smooth."
Cloudflare, 20 February. A cleanup task called their own API with pending_delete and no value after it.
The API read that as: everything. 1,100 of 4,306 prefixes withdrawn from BGP. https://t.co/tJB40UMByv served a 403.
Six hours, from one missing value in a query string.
Error tracking can only report what your code observed. It cannot report a request that never arrived.
Supabase, 12 February: a config change blocked all internet gateway traffic to us-east-2. No errors were logged, because there were no requests.
Watch from outside.