Monitoring that tells you something: status pages and alerts worth having

Most monitoring produces either silence or noise. What a useful status page measures, why percentiles beat averages, and how to build alerting that someone actually acts on.

Most organisations have monitoring. A lot of it sits in one of two states: a dashboard nobody looks at, or an alert channel so noisy that everyone has muted it. Both fail in the same way. When something breaks, a customer notices first. Useful monitoring answers a small number of questions reliably: is the service working for the people who use it, how do we know, and who finds out when it is not?

Measure what users feel first

There are two kinds of measurement, and they do different jobs. Symptoms are what users experience: whether the service responds, how long it takes, and whether the answer is correct. Causes are what is happening inside: CPU, memory, disk space, queue depths, replication state.

Symptoms belong at the top. A server running at 95 per cent CPU while every request completes quickly is not an incident. A server idling at 20 per cent while every request times out is. Cause metrics are essential for working out why something is wrong, but they make poor primary alerts, because they fire when nothing is wrong and stay quiet when something is.

The most direct way to measure symptoms is a synthetic check: a probe that makes a real request on a schedule, the way a user would. A full HTTPS request to the login page, a call to the API, a test call through the phone system. Frequency matters. A check every minute gives useful resolution. A check every hour can miss an outage that lasts most of an hour.

Averages hide the pain

Average response time is the most commonly reported number and one of the least useful. It is dominated by the typical request, and the slow ones that frustrate people disappear into it.

Percentiles fix this. The 95th percentile, or p95, is the time that 95 per cent of requests beat, which means one request in twenty is slower. If a page averages 200 milliseconds but its p95 is three seconds, one visit in twenty is painful, and a staff member who loads that page fifty times a day meets the slow tail every day. For busy services the 99th percentile is worth watching as well.

Percentiles need enough samples to mean anything. One-minute checks over 24 hours give 1,440 samples, enough for a stable p95. A p95 calculated over the last ten checks is mostly noise.

Separate the network from the server

A response time measured from outside includes several things: resolving the name, opening the connection, negotiating TLS, the server doing its work, and the response travelling back. A slow total might mean a slow application or a congested link, and those have different owners.

Breaking the number apart helps. Time to first byte, minus the connection and TLS setup, approximates the server’s own work. Network round-trip time and TLS handshake time, reported separately, show what the path is contributing. When the server time is steady and the round trip has doubled, the conversation is with the network provider, not the developers.

This is how the Bizix status page is laid out. Every minute, from Sydney, it records server response time as time to first byte minus TLS setup over HTTPS, the p95 of those one-minute samples over the last 24 hours, network round trip and TLS handshake time for context, and uptime over the last 24 hours.

What a status page is for

A status page answers the first question anyone asks during an incident: is it us, or is it them? Answered quickly and honestly, it saves a flood of support calls and a lot of guesswork on both sides.

Honesty is the whole value. A page that only ever shows green is decoration. A useful one shows degraded states as well as outages, gives times, is updated during an incident rather than after it, and keeps enough history that people can see the pattern.

It should also be clear about its scope. Our page separates the services we operate from the upstream providers shown for context, and the headline status reflects only the services we run. A provider incident may explain something a user is seeing without our infrastructure being at fault, and saying so plainly is more useful than a single traffic light. The page also states that it is our own vantage point, not a substitute for independent external monitoring. A single vantage point is a start, not a finish, and any status page should say which one it is.

Alerts are for humans

An alert is a request for a person to do something now. If nothing needs doing now, it is not an alert. It is a ticket, a report or a line on a dashboard.

  • Every alert should be actionable and urgent, and should reach a named person on call rather than a shared inbox.
  • Alert on symptoms, with a sensible duration: three failed checks in a row, not one, so that a single dropped packet does not wake anyone.
  • Every alert should link to a short runbook covering what it means, what to check first and who to escalate to.
  • Trends belong in business hours. A disk forecast to fill in ten days is a ticket. A disk forecast to fill in two hours is a page.
  • On a hyperconverged cluster, a failed disk that the cluster is already rebuilding around is a ticket. A cluster that has lost redundancy and cannot rebuild is a page.

Then review. After every incident, ask whether monitoring caught it before a user did. After every noisy week, fix or delete the alerts that fired without anyone needing to act. The number of pages per on-call shift is itself a metric worth watching.

Watch the watcher

Monitoring fails quietly. The probe host dies, a credential expires, the mail relay that delivers alerts breaks, or the on-call number changes and nobody updates it. From the inside, a broken monitoring system and a healthy environment look identical: nothing is alerting.

Two habits cover most of this. First, a heartbeat: the monitoring system sends a regular signal to something independent of it, which raises an alarm if the signal stops. Second, end-to-end tests: on a schedule, deliberately trigger a known failure and confirm that a real person receives the alert on the device they actually carry.

A practical starting point

  • List the five services that matter most, and give each a user-facing check that runs every minute.
  • Record p95, not just the average.
  • Delete or demote any alert that nobody acted on last month.
  • Add a heartbeat for the monitoring itself, and test the alert path end to end.
  • Publish a status page that is honest about what it measures and where from.

Where Bizix fits

Bizix publishes its own live status page, checked every minute from Sydney, and telemetry and monitoring integration are part of the hardening phase of every Bizix private cloud deployment. The environments we operate are covered by 24/7 incident response from the team that designed them, so the person who receives the alert is someone who knows what it means.