The checkout dashboard has a panel called payment_gateway_queue_depth. It has been at 0 for eleven months. Nobody noticed, because nobody looks at it. It was added at 2am during an incident, and the person who added it left the team in March.
That is the normal story. A panel appears when we are scared, it answers one question once, and then it stays on the dashboard forever. Within a year it is flat, wrong, or so noisy that people scroll past it.
Does knowing it change what you do at 3am?
This is the only test that matters. Take a panel and ask: if this number were alarming right now, what would I do differently? If the answer is "nothing, but it would be interesting," the panel is decoration.
I use a blunter version with my team. Pick the metric. Now finish this sentence: "If this spikes, I will ..." If you cannot fill in the blank with a command, a rollback, or a page to another team, the metric is not operational. It is trivia with a graph attached.
The same test deletes alerts. An alert that fires and gets ignored is worse than no alert, because it trains everyone to treat the pager as background noise. If an alert has fired more than a handful of times and nobody acted on any of them, it is a deletion candidate. Either fix the threshold so it only fires when action is needed, or delete the rule and keep the graph for humans to read during incidents.
Collected because it was easy
Most dead metrics got there the same way: the exporter emitted it for free, so we scraped it. Request rate per endpoint, response size percentiles, cache hit ratio by region — all collected, none attached to a decision.
There is a real difference between a metric with a decision attached and one collected because it was easy. The first one has a threshold, an owner, and a runbook. The second one has a dashboard slot.
Service-level indicators sit on the decision side. They describe what the user experiences: successful checkouts per minute, search latency under 500ms, payment authorization success rate. Vanity throughput graphs sit on the other side. Total requests served tells you the system is busy. It does not tell you whether anyone got what they came for.
I have watched a service serve 40% more requests than the week before while its success rate dropped, and the throughput graph looked like a win. The SLI panel next to it was the one that mattered.
When a service has no SLI, that is the panel to build. Not another throughput chart.
Attach an owner and a line of intent
Every panel gets two pieces of metadata, in the dashboard config itself, not in a wiki page nobody reads.
panels:
- title: payment_gateway_queue_depth
owner: payments-team
intent: "If depth > 500 for 10m, page payments on-call and check gateway health."
query: sum by (gateway) (payment_gateway_queue_depth)
alert: PaymentGatewayQueueBacklog
last_reviewed: 2026-09-14
The owner field is a real team, not a person, because people rotate. The intent line is the "what would we do about this" sentence, written down. If nobody can write that line, the panel does not get created.
last_reviewed is the part that keeps the whole thing alive. Once a quarter, walk the dashboard and read those dates. Anything older than two quarters gets either a fresh intent line from its owner or a delete commit. A stale panel is a claim that someone is watching. That claim is false, and it is worse than an empty dashboard because it makes us believe we have coverage.
The graph that is flat at 0 is the dangerous one. It looks healthy. It looks calm. It has not received a single data point since the exporter changed its metric name in the spring.
Measure the decay instead of guessing
Do not estimate how many panels are dead. Measure it. For each panel, query the underlying series over the last 30 days and count series with zero variance. For each alert rule, count firings and correlate them with the incident channel in the same window. Alerts that fire with no corresponding human action are your deletion list.
Then delete. Not archive — delete, so the next person who opens the dashboard sees only things that have an owner and a reason.
A dashboard is a set of promises that someone is watching. Keep the promises you can keep, and remove the rest. The panel that nobody owns does not monitor anything. It just makes the silence look intentional.
I write about production failures in Postgres, queues, and distributed systems.
Top comments (0)