- Blog
- Observability
- APM versus traditional monitoring: what each one sees, and when you need both
Observability
APM versus traditional monitoring: what each one sees, and when you need both
APM and infrastructure monitoring answer different questions. What each sees, what each misses, and which your estate needs first.

Two teams can look at the same outage and describe it in unrelated words. The infrastructure team says the database host hit 95 % CPU at 14:02. The application team says checkout latency went from 300 ms to 9 s because one query started doing a sequential scan. Both are right, and neither saw what the other saw, because they were using different kinds of monitoring. This article is about what those kinds are, what each can and cannot see, and how to decide what your estate needs.
The short version: traditional monitoring watches the things applications run on; application performance monitoring (APM) watches the requests running through them. Most estates need the first everywhere and the second in a few places, and getting the order right saves a great deal of money and attention.
What traditional monitoring sees
Traditional monitoring — infrastructure monitoring, if you prefer — measures resources and availability from the outside. For a host: CPU, memory, disk, network. For a database: connections, locks, replication lag, storage growth. For an endpoint: did it answer, how quickly, with what status. For a scheduled job: did it start, did it finish, how late. For an integration platform: workers, throughput, errors, queue depth.
Its strengths are breadth and simplicity. Every kind of system in an estate can be covered the same way, the measurements are cheap to collect, the rules are easy to explain, and a new system can be connected in minutes. It is the layer that tells you a disk will be full on Thursday, a job did not run, a collector went quiet, a certificate expires next week.
Its blind spot is the inside of a request. When the checkout endpoint takes 9 s, infrastructure monitoring can tell you the host is busy and the database is under load. It cannot tell you which of the twelve calls the endpoint makes is the slow one, or that the slow one is slow only for customers with more than a hundred items in the basket.
What APM sees
APM instruments the application itself. An agent in the process records each request as a trace: the endpoint, the time spent in each function, each outbound call to a database, a cache or another service, and how long each took. Traces are sampled, aggregated and searchable, so that a question like "which endpoint got slower after the deploy at 11:40" has an answer.
Its strength is depth. It turns "checkout is slow" into "the PriceBasket call is spending 8 s in a query against inventory that used to take 40 ms", and it does so for the specific request a customer complained about. For a team whose recurring problem is performance regressions in code they own, nothing else answers the question.
Its blind spots are everything the agent is not inside. A host that is out of disk, a scheduled job that did not start, a certificate that expired, an integration platform whose workers have stalled, a database whose replication has fallen behind: APM sees none of these directly. It sees their consequences, sometimes, once requests start failing, and by then infrastructure monitoring would have raised a warning an hour earlier.
There is a second cost, less often mentioned. APM asks the team to learn a new way of reading: traces, spans, flame graphs, percentiles. That is a real skill, and until it is acquired, the data is a wall of detail that nobody uses during an incident.
Side by side
| Question | Traditional monitoring | APM |
|---|---|---|
| Is the host running out of disk? | Yes, and projects when | No |
| Did the nightly job run? | Yes | No |
| Is the endpoint answering? | Yes, from outside | Yes, from inside |
| Which function is slow? | No | Yes |
| Which query is slow? | Sometimes, from the database's side | Yes, from the request's side |
| Is the integration platform's queue backing up? | Yes, if it is a first-class subject | Rarely |
| Did latency change after a deploy? | Coarsely, per endpoint | Precisely, per span |
| Effort to cover a new system | Minutes | Hours to days, per application |
| Coverage of things you did not write | Complete | None |
The last two rows decide the order for most estates. Infrastructure monitoring covers everything, including the systems you bought rather than built, for a small effort each. APM covers only what you can instrument, for a larger effort each, and answers questions that only some teams are asking.
A typical estate, and what each layer would have caught
Take an estate with a customer-facing web application, its database, an integration platform moving orders to a warehouse system, a CRM with an API, and three scheduled jobs. Over a quarter, the incidents that occur might look like this.
- Database disk fills over three weeks and the application starts failing writes. Infrastructure monitoring projects the exhaustion a week out. APM sees the write failures on the day.
- The nightly stock load stops being scheduled after a host rebuild. Infrastructure monitoring's "did not start by 02:15" rule fires that night. APM never sees it; no request touched the job.
- Checkout latency rises from 300 ms to 4 s after a release. Infrastructure monitoring sees a slower endpoint and a busier database. APM identifies the query and the code path in minutes.
- The integration platform's workers stall and orders queue for two hours. Infrastructure monitoring, with the platform as a first-class subject, alerts on queue depth in ten minutes. APM sees nothing until a customer asks where their order is.
- The CRM's API rate limit is exhausted by a new sync and contacts go stale. Infrastructure monitoring watching the API's own limit counters catches it. APM might see slow outbound calls if the sync is instrumented.
- A memory leak in one service grows over four days. Infrastructure monitoring projects the exhaustion. APM, if it tracks heap per service, sees it too; if it does not, it sees the eventual restarts.
Four of six caught earlier by the infrastructure layer, one by APM, one by both. That ratio is typical, and it is why the infrastructure layer comes first. The one that only APM catches — the regression in your own code — is also the one where the difference between "the endpoint is slow" and "this query is slow" is worth the most, which is why APM comes second rather than never.
Deciding for your estate
Three questions settle it.
How much of the estate did you write? If most incidents are in systems you operate but did not build — databases, integration platforms, SaaS, batch jobs — APM has little to instrument and infrastructure monitoring is the whole answer. If most incidents are performance regressions in your own services, APM earns its place.
Can anyone currently answer "which part is slow"? If the answer is a shrug and a guess, and the question comes up monthly, that is the case for APM on the two or three services where it comes up. Instrument those, not everything.
Who will read it? APM data is read by developers, during and after incidents, and it rewards practice. If the on-call rotation is an operations team that does not own the code, APM will be a screen they open once. Give them the infrastructure layer with a service-level signal per application instead, and give the developers APM for the services they own.
Doing both without doubling the noise
Estates that run both layers well follow a few rules.
One alert per failure, from the layer that sees it first. Do not alert on high latency from both APM and the endpoint check; pick the one that says more. Route infrastructure alerts to operations and APM alerts to the owning developers.
Link across, do not merge. A useful alert from the infrastructure layer names the service and links to its traces; a useful trace view shows the host's resource state alongside. The two layers stay separate in what they collect and join in what they show.
Instrument by incident, not by inventory. Add APM to a service when an incident in it went unanswered, not because it exists. A short list of instrumented services that people actually consult beats full coverage that nobody reads.
Keep the service-level signal regardless. Even with APM on a service, keep the outside-in check: does the endpoint answer, how fast, with the right shape. An agent inside a process that has stopped accepting connections reports nothing, and the outside check is what notices.
In short
Traditional monitoring sees the estate; APM sees the request. Start with the first, everywhere, with a signal per service that the service itself produces. Add the second where the question "which part is slow" is asked and unanswered, to the services whose owners will read the answer. Route each layer's alerts to the people who can act on them, and let each link to the other rather than repeating it.
That order covers the incidents that actually happen, in the order they actually cost, and it leaves the team with a set of screens they open rather than a set they were sold.