- Blog
- Monitoring
- Real-time infrastructure monitoring: practices that survive contact with production
Monitoring
Real-time infrastructure monitoring: practices that survive contact with production
What real-time means in infrastructure monitoring, and the collection, threshold and routing practices that keep alerts honest at 3 am.

"Real-time" is the most overloaded phrase in monitoring. To a vendor it means a chart that updates while you watch. To an engineer on call it means one thing only: when a system goes wrong, how long before someone who can fix it knows. Measured that way, most estates are not real-time at all. They are "within an hour, if someone is looking", and the gap between the two is where outages become incidents.
This article is about closing that gap without drowning the team. It covers what to collect and how often, how to write thresholds that fire on change rather than on noise, how to route alerts so they are read, and the handful of practices that separate a monitoring setup that is trusted from one that is tolerated.
What real-time means, measured
Three delays add up to the time between a failure and a person acting on it.
- Collection delay: how long between the condition and the platform holding a measurement of it. Set by the collection interval.
- Evaluation delay: how long between the measurement arriving and a rule deciding it matters. Set by the rule's "held for" window.
- Delivery delay: how long between the rule firing and the right person seeing it. Set by the channel, the routing and the time of day.
A sensible target for a customer-facing system is five minutes end to end during working hours and fifteen outside them. Faster is possible; it is rarely worth what it costs in false alarms. Write the target down per system, because the nightly batch job and the checkout API do not deserve the same one.
Collect the right things at the right rate
The four signals per host
For a host or virtual machine, four measurements answer almost every question: CPU utilisation, memory used, disk used and network throughput. Add disk I/O wait if the host runs a database, and the process count if it runs something that forks. Everything else is detail you look up during an incident, not something you watch continuously.
Collect these every 30 to 60 seconds. One-second collection produces beautiful graphs and 60 times the storage, and no decision changes because of it. Five-minute collection is common and too slow: a spike that lasts three minutes never appears, and a memory leak looks smooth right up to the moment it is not.
The service's own signal
A host can be healthy while the service on it is dead. For each service that matters, watch a signal the service itself produces: an HTTP endpoint that returns 200 with a body, a queue depth, a "last processed" timestamp, a row count that should grow. This is the measurement that tells you the service is doing its job, and it is the one most estates lack.
For an endpoint, measure three things: whether it answered, how long it took, and whether the answer was the right shape. A 200 in 40 ms with an empty body is a failure that a simple up/down check calls healthy.
Scheduled work
Anything that should happen on a schedule needs two measurements: when it last started and when it last finished. From those two, and the schedule, the platform can tell you that a job is late, is running long, or did not happen at all. The last one is the failure that a job's own error handling can never report, because there is no job running to report it.
What not to collect
Do not collect what you would not act on. Every extra series is storage, a slower query and one more line on a graph that hides the one that matters. When somebody proposes a new metric, ask what alert rule or decision it feeds. If the answer is "it might be interesting", leave it out until it is.
Thresholds that fire on change
A threshold is a claim: above this value, something is wrong. The claim is only true relative to what normal looks like for that system, which is why copied thresholds fail.
Learn normal first
Collect for at least a week before writing a rule. Look at the daily cycle and the weekly one. A system that idles at 20 % CPU and peaks at 55 % on Monday mornings has a very different "wrong" from one that runs at 70 % all day. The first should alert at 75 %; the second should alert on a change from its own baseline.
Hold the condition
A rule that fires the instant a measurement crosses a line fires on every spike. Require the condition to hold: CPU above 80 % for ten minutes, disk above 90 % at three consecutive readings, endpoint latency above two seconds for five minutes. The hold window is the single most effective noise reduction available and it costs nothing but a little delay you have already budgeted for.
Project, do not just compare
Disk full is a threshold. Disk full on Thursday is a projection, and it is the one that lets a team act during the day. Where the platform can project a metric from its trend, alert on the projection crossing the threshold within a window — "CPU reaches 80 % within two hours" — and say what the projection is based on, so that a reader can decide whether four minutes of trend is enough to believe.
Treat silence as a condition
The most important rule in any estate is the one that fires when a system stops reporting. It catches the collector that died, the host that was decommissioned without telling anyone, the network path that closed. Write it for every system, with a window that matches the collection interval — no data for three intervals is silence, one missed interval is a hiccup.
Make sure the platform shows silence differently from health. A flat line at zero and an empty chart mean opposite things, and a screen that draws them the same way will be misread.
Route alerts so they are read
One system, one owner
Every alert needs a named recipient, and the recipient should be the person or team who can act. Alerts to a shared inbox that six people are meant to watch are alerts nobody watches. Record the owner when the system is connected and route by that record.
Severity is about action, not feeling
Use two levels, three at most. Critical means somebody should stop what they are doing, including sleeping. Warning means somebody should look during the next working period. A third level, if you need one, is notice: recorded, visible, never delivered. Anything that does not fit one of these is not an alert; it is a metric that belongs on a dashboard.
The test for critical is simple: would you be right to wake someone for this? If the honest answer is "it can wait until morning", it is a warning.
Prove the channel
Every channel — email, chat, phone, push — should be tested with a real message when it is set up and again whenever it changes. Email is the channel most likely to fail silently: the recipient's mail system quarantines it, and neither side hears. A platform that records what became of every message it sent, and that lets you send a test to yourself in one click, turns "I never got the alert" from an argument into a lookup.
Say enough to act
An alert that reads "CPU high on host-14" sends the recipient to find out what host-14 is, what it runs and what high means. An alert that reads "Orders API host: CPU 91 % for 12 min (usual peak 55 %); at this rate memory reaches 85 % in 40 min; owner: platform team" can be acted on from the phone. Include the system's name in plain words, the measurement and its normal, the trend if there is one, and the owner.
The practices that survive
The techniques above are known. What separates estates where they are still in force a year later is a small number of habits.
Review the alerts, not the dashboards. Once a month, list every alert that fired. Mark each as acted on, acknowledged-but-nothing-to-do, or ignored. Retire or retune anything ignored twice. Thirty minutes of this keeps the signal readable; skipping it is how estates reach the point where nobody reads anything.
Break it on purpose. Once a quarter, stop a collector, pause a job, block a channel. Confirm the platform noticed, said so, and reached the right person. A monitoring setup that has never been tested against absence is a hypothesis.
Keep the reason with the rule. Why is the threshold 85 %? Because at 88 % on the 14th the queue backed up. Write that in the rule's description, where the person changing it at 3 am will read it.
Watch the watcher. The platform needs a rule for its own silence, delivered by a path that does not depend on the platform. If the collectors stop, somebody should know within the same five minutes you promised for everything else.
Prefer a plain sentence to a clever chart. The most useful screen in an estate is one that says "nothing is at risk right now; the next projected breach is the payments API's CPU in about four hours". Anyone can read it. A wall of forty sparklines is read by nobody.
A short checklist
For each system that matters:
- Collected every 30–60 s, including a signal the service itself produces.
- One rule for a real failure mode, held for a window, with the reason written on it.
- One rule for silence.
- One named owner, with a proven channel.
- A severity that answers "should we wake someone?".
For the estate:
- A monthly alert review with a retire-or-retune rule.
- A quarterly deliberate break.
- A rule for the platform's own silence.
- A screen that answers "is anything at risk?" in one sentence.
An estate that meets this list is real-time in the sense that matters. The charts updating while you watch are a bonus.