- Blog
- Operations
- The return on a monitoring platform: counting downtime you did not have
Operations
The return on a monitoring platform: counting downtime you did not have
A model for what a monitoring platform is worth: the cost of an hour down, the minutes removed per incident, the incidents prevented, and the false alarms.

A monitoring platform is bought on a promise about downtime and renewed on a feeling. Neither is a number. This article is about producing one: a model you can fill in for your own estate that says what an hour down costs, how many minutes the platform removes from each incident, how many incidents it prevents, and what it costs in return — including the cost of the alerts nobody should have received.
The figures in the worked example are invented for the example. Resist the urge to borrow them; the whole point is that they are yours to fill in.
The three parts of the cost of downtime
An hour of downtime on a system costs three things, and estates that quote one number for all systems have usually only counted the first.
Revenue that does not happen. For a system that takes orders or bookings, the sales that would have occurred in the window. Some return when the system does; some go elsewhere. A useful approximation is the average hourly revenue for that hour of that day, times the share that does not come back — a share you can estimate from the last real outage if you had one.
Work that has to be redone or done by hand. A failed nightly load means a morning of manual reconciliation. A stalled integration means orders keyed by hand or held until it recovers. Count the hours of people's time at their loaded cost, and add any penalty the delay triggers with a customer or a partner.
People who stop. Every incident pulls in more than the person fixing it: the manager asking for updates, the support team fielding calls, the colleague whose report is late. Count the number of people and the hours they lose. This is usually the largest part for internal systems and the most often omitted.
Work this out per system, per hour of the day, at least for the ten systems that matter. The result is a table like this (invented figures):
| System | Hour | Revenue lost per hour | Rework per hour | People stopped | Cost per hour |
|---|---|---|---|---|---|
| Checkout API | 14:00 weekday | 4,200 | 300 | 6 × 60 = 360 | 4,860 |
| Checkout API | 03:00 | 180 | 0 | 1 × 60 = 60 | 240 |
| Nightly stock load | 02:00 | 0 | 900 (morning reconciliation) | 4 × 60 = 240 | 1,140 |
| CRM sync | 10:00 | 0 | 150 | 8 × 20 = 160 | 310 |
Two things fall out immediately. The same system is worth twenty times more at 14:00 than at 03:00, which is why alert routing by time matters. And the batch job that has no revenue attached costs a quarter of the checkout API per hour, which is more than most teams assume.
What a platform changes: detection delay
The clearest effect of a monitoring platform is on the time between a condition beginning and somebody knowing. Call it detection delay. Before a platform, for many estates, it is "when a customer calls" or "when somebody looks at the report", which is measured in hours. After, for a well-configured platform, it is the collection interval plus the rule's hold window plus delivery: five to fifteen minutes.
Detection delay is the part of every incident's duration that the platform removes outright. Resolution time — how long it takes to fix once known — is affected too, by better context in the alert, but that effect is smaller and harder to attribute. Be conservative and count detection alone.
To measure it, record three timestamps for every incident: when the condition began (from the data, after the fact), when a person knew, and when it was resolved. Most teams find they can reconstruct the first from the platform's own history once it exists, which is itself one of the platform's quieter returns.
What a platform changes: incidents that do not happen
The second effect is prevention: conditions caught as trends and fixed during working hours. A disk projected to fill on Thursday is extended on Tuesday. A memory leak projected to exhaust a host in four days is restarted at lunchtime. A batch job drifting later each night is fixed before it crosses the report deadline.
These are harder to count because the incident that did not occur leaves no record. Count them anyway, from the platform's history: every projection-based alert that was acted on before the threshold was crossed is a prevented incident, and its avoided cost is the cost per hour of that system times a conservative estimate of how long the resulting outage would have lasted — an hour is a defensible floor for anything that would have needed a person to notice and intervene.
The worked model
Here is the model with invented figures for an estate of forty systems, ten of which matter, over one quarter.
Before the platform (reconstructed from the previous quarter's incident log):
- 14 incidents on the ten systems that matter.
- Average detection delay: 74 minutes (found by customers, by a morning check, or by luck).
- Average cost per hour across those incidents, weighted by system and hour: 1,900.
- Detection cost per quarter: 14 × 74 min × 1,900 / 60 = 32,800.
After the platform (first full quarter):
- 11 incidents (three fewer; the disk, the leak and the drifting job were caught as trends).
- Average detection delay: 9 minutes.
- Detection cost: 11 × 9 min × 1,900 / 60 = 3,100.
- Prevented incidents: 3, at a conservative one hour each at 1,900 = 5,700 avoided.
Gross return per quarter: (32,800 − 3,100) + 5,700 = 35,400.
Costs per quarter:
- Licence: 2,400.
- Time keeping it configured: 3 hours a week × 13 weeks × 60 = 2,340.
- False alarms: 22 alerts investigated and dismissed × 20 minutes × 60 / 60 = 440.
- Total: 5,180.
Net return per quarter: 30,200, or roughly six times the cost. Change any figure and the answer changes; the shape of the calculation does not.
The false-alarm line deserves a note. In this example it is small because the estate followed the discipline of one rule per real failure mode and a monthly review. An estate with three hundred template rules can see that line reach thousands, and a platform whose alerts are ignored has a detection delay of "when a customer calls" again, whatever the dashboard says.
The returns that do not fit the model
A few effects are real and resist a number. List them in the business case as what they are rather than inventing a figure.
Incidents explained rather than argued. A recorded history of every measurement and every alert message ends the meeting where two teams disagree about what happened at 14:02.
Change made with evidence. Capacity is added when a projection says so, not when a manager feels nervous or a host falls over.
On-call that people will do. A quiet, honest alert stream is the difference between a rota that engineers accept and one they leave over.
Customers told first. When the platform knows before the customer does, the customer hears from you rather than the reverse. That is worth more than the revenue line in most relationships.
Measuring it after the fact
The model is a forecast. The proof is the same three timestamps per incident, kept for every quarter, and a simple report:
- Number of incidents on the systems that matter.
- Median and average detection delay.
- Incidents caught as trends and fixed before they occurred.
- Alerts investigated and dismissed.
- Cost per hour table, revised as the estate changes.
If detection delay is falling, prevented incidents are rising and dismissed alerts are flat or falling, the platform is earning its keep. If dismissed alerts are climbing, the platform is not the problem; the rules are, and the monthly review is where that gets fixed.
Filling it in for your estate
The whole model needs six inputs you can gather in a week:
- The ten systems that matter, from the estate table.
- Cost per hour for each, at the hours that matter, from the three parts above.
- Last quarter's incidents on those systems, with detection delay reconstructed as best you can.
- The platform's licence and the honest hours of attention it will need.
- A target detection delay, per system, from the acceptable-delay column of your estate table.
- A rule for counting prevented incidents: one hour of that system's cost per trend caught before it crossed.
Put those in the model and you have a number that will survive a finance review, because every part of it is either measured or labelled as an estimate with its basis. That is a better footing than a quoted industry average, which describes an estate that is not yours, at an hour that is not the one that hurts.