Monitoring a client can read
The dashboard we hand over after launch, what is on it, and why error counts and uptime charts are not on the first screen.
· 5 min read
Every monitoring tool ships with a default dashboard, and every default dashboard is built for the person who installed it. Request rates, p99 latency, error counts, memory. All useful. None of it answers the question a business owner actually has, which is: is my product working for the people paying for it?
So we build two views. The engineers keep the dense one. The client gets one that can be read in fifteen seconds, on a phone, by someone who does not know what a p99 is.
The first screen answers one question
At the top there is a single line in plain language: everything is working, or something is not and here is what. Not a green dot — a sentence. A green dot still requires the reader to trust that the dot is measuring the right thing.
Underneath it are the two or three things the business actually cares about, expressed in its own vocabulary. For a store: orders in the last hour, against the same hour last week. For a booking product: bookings completed, and bookings abandoned at payment. For an internal tool: tasks processed, and tasks waiting for a human.
If a metric would not change what anyone does on Monday, it does not belong on the first screen.
Why error counts come second
Error counts are a poor headline for two reasons. They are almost never zero — a healthy system has a steady trickle of bots, expired tokens and people closing tabs mid-request — so a non-zero count alarms the reader without telling them anything. And a count treats every error alike, when one failed checkout matters more than four hundred blocked scrapers.
We keep errors on the second screen, grouped by what broke rather than by how often, and we set alerts on the specific failures that mean revenue is being lost. Those alerts go to a person who can act, not to a shared inbox where they will be admired.
Uptime is a contract term, not a dashboard
A 99.95% uptime chart is a reporting artefact. Nobody makes a decision from it. What matters is whether the product was usable during the hours the business trades, and whether anyone had to be told. So the client view shows incidents — when, how long, what we did — and leaves the percentage to the monthly report.
What we send every month
- What changed: shipped features and fixes, in plain sentences.
- What broke: every incident, its cause, and what we changed so it does not recur.
- What it cost: infrastructure spend, and anything trending in a direction worth discussing.
- What we would do next, with a recommendation rather than a menu.
The last one matters most. A monthly report that only looks backwards makes the client responsible for deciding what happens next, which is the job they hired us to have an opinion about.
The underlying instrumentation is the ordinary kind — structured logs, traces, error tracking, alerting. The difference is only that somebody sat down and asked what the reader needs to know. That is an hour of work, and it is the difference between a dashboard that gets opened and one that gets bookmarked and forgotten.