- Blog
- Observability
- Log aggregation and analysis: turning logs into something you can ask questions of
Observability
Log aggregation and analysis: turning logs into something you can ask questions of
Why logs on twelve hosts are not observability: what to ship, what to index, how to carry a trace id across services, and what to alert on.

Every estate has logs. Very few can answer a question with them. The difference is not the volume — a busy estate produces more log lines in an hour than anyone will read in a year — but whether the lines can be found, related and trusted when a question arrives. This article is about the practices that make that possible: what to collect, how to shape it, what to index, how to follow one request across many systems, and what is worth alerting on.
It assumes a mixed estate: a few applications you wrote, some you bought, a database or two, an integration platform, scheduled jobs. That mix is why log aggregation is harder than a single-application tutorial suggests, and why the practices below favour honesty about gaps over pretending every source is the same.
Why logs on twelve hosts are not observability
Logs left where they are written answer one question: what did this process say. During an incident, the question is different: what happened to this request, or this customer, or this order, across everything it touched. Answering it from twelve hosts means twelve terminals, twelve time zones of clock drift, twelve formats, and a person doing the joining by hand at the worst possible moment.
Aggregation moves the lines to one place with one clock. That is necessary and not sufficient. Observability — the ability to ask a question you did not plan for — needs three more things: structure, so the lines can be filtered; a shared identifier, so they can be related; and retention, so the question that arrives a week late can still be answered.
Collect: ship everything, index deliberately
Ship all of it
Do not filter at the source. The line you drop today is the one you need next month, and the cost of shipping a line is far below the cost of the incident it would have explained. Ship every application log, every job log, the database's slow-query log, the integration platform's runtime log, the web server's access log and the operating system's authentication log.
Ship them with the host, the service and the environment attached as fields. A line without those is a sentence without a speaker.
Index a little
Shipping everything and indexing everything are different decisions. Full-text indexing every field on every line is what makes log platforms expensive. Index the fields you will filter on — time, host, service, level, trace id, and a handful of domain identifiers such as order id or customer id — and leave the message body as searchable text. That keeps queries fast and the bill readable.
Keep the failures visible
Lines that could not be parsed, sources that stopped sending, clocks that drifted: record each as a condition rather than dropping it silently. A source that goes quiet should show as absent, not as a host that has simply stopped having problems. The most dangerous state a log platform can be in is one where a collector has died and the dashboard reads clean.
Shape: structure at the source
Structured lines
An application you control should write structured lines: a timestamp, a level, a message, and named fields. JSON is the common choice because every platform reads it. The rule is that a machine reads the line before a person does, so the machine's needs come first: consistent field names, one event per line, no multi-line stack traces spread across five lines that a collector will split.
{"ts":"2026-09-17T14:02:11.412Z","level":"warn","service":"orders-api",
"trace_id":"7c1e…","order_id":"A-88213","msg":"inventory lookup slow",
"duration_ms":8410,"query":"select … from inventory where …"}
Everything in that line is filterable. The equivalent prose line — "WARN inventory lookup for order A-88213 took 8410ms" — has to be parsed with a pattern that breaks the day somebody rewords it.
Parse the rest at ingestion
Systems you did not write will log in whatever shape they log in. Parse those lines at ingestion with a pattern per source, extract the same core fields, and keep the original line intact alongside. When a pattern stops matching — after an upgrade, typically — record the miss as a metric and alert on it, because a source whose lines have all become "unparsed" has effectively gone dark.
Levels that mean something
Four levels are enough: error (a request or job failed), warn (something is degraded or unexpected but the work completed), info (a notable event: started, finished, deployed), and debug (off in production). The discipline that matters is that "error" is reserved for failures a person would want to know about. An estate where errors are logged for expected conditions cannot alert on error rate, which is the single most useful log-derived alert there is.
Relate: the trace id
Mint once, carry everywhere
When a request enters the estate — at the web server, the API gateway, the integration platform's inbound endpoint — mint an identifier and attach it. Every service that handles the request logs the identifier on every line and passes it on every outbound call, in a header for HTTP, in a message property for queues, in a column for jobs that pick up work later.
That one practice turns a search from "find lines about order A-88213 across six systems and guess which belong together" into "show me trace 7c1e…", returning the whole story in time order. It is the difference between an hour and a minute, and it costs one header.
Where it breaks, and what to do
The trace id breaks at every boundary you do not control: a SaaS application that will not echo your header, a batch job that reads from a file, a vendor integration that mints its own. At each, log both identifiers on the same line — yours and theirs — so the join is recorded rather than lost. A line that says "handing order A-88213 (trace 7c1e…) to CRM as record 00Q5g…" is the bridge a search will follow later.
Clocks
Related lines are only related in time if the clocks agree. Synchronise every host and every container, log in UTC, and record the collector's receive time alongside the line's own timestamp. When the two disagree by more than a few seconds, that is a condition worth surfacing: a host with a drifting clock is a host whose lines will sort into the wrong place in every investigation.
Analyse: the questions worth being able to ask
A log platform earns its keep by answering a small number of questions quickly. Make sure each of these is a saved query, not a skill.
- What happened to this request? By trace id, across every service, in time order.
- What is failing right now? Error lines in the last fifteen minutes, grouped by service and message, sorted by count.
- What changed? Error rate per service, this hour against the same hour last week, and the deploy events in between.
- Who did this? Authentication and audit lines for a user or an API key over a window.
- How slow, and for whom? Duration fields, by endpoint, at the 95th percentile, split by any identifier that might explain the tail — customer, region, basket size.
- What did the job do? Every line for a job run, from start to finish, with the counts it reported.
If a question in this list takes more than a minute to answer, the fix is usually a missing field at the source, not a faster search engine.
Alert: from logs, sparingly
Logs make poor alert sources in general — irregular, verbose, and expensive to scan — and excellent ones in a few specific cases.
Error rate per service. Count error-level lines per minute per service and alert when the rate exceeds its own recent baseline. This catches failures that never cross a resource threshold: a downstream returning 500s, a bad deploy, a configuration change. It depends entirely on the level discipline above.
A line that should never appear. An unhandled exception, a failed authentication burst, a "deadlock detected", a certificate warning. Write a rule per message pattern, with the reason on the rule.
A source that went quiet. No lines from a service that normally writes every second is an outage or a dead collector; either way somebody should know.
A parse pattern that stopped matching. The upgrade-broke-the-format condition, caught the same day rather than during the next incident.
Everything else — latency, throughput, saturation — belongs to metrics, which are made for regular evaluation. Use logs to find out what happened, not to notice that it did.
Retention and cost, honestly
Storage is the cost that grows quietly. A workable policy is three tiers: the last thirty days indexed and fast; thirty to ninety days searchable but slower; older than ninety days in cheap object storage that can be reloaded on request. Delete nothing before a year unless the data is personal and a policy requires it, in which case delete it on schedule and record that you did.
Review what is being shipped once a quarter. The source producing half the volume is usually a debug flag left on or a health check logging every poll; turning it down costs nothing and pays every month.
Where this meets monitoring
Logs are one of three signals, alongside metrics and traces, and they are the most detailed and the least regular. A monitoring platform tells you that the orders API's error rate doubled at 14:02; the log platform tells you it was the inventory query, for baskets over a hundred items, after the 11:40 deploy. The two are most useful when the alert from the first carries the trace id or the service name that opens the second at the right place.
If your monitoring platform can follow a trace id across the estate itself — showing one request's path with the lines each service wrote — the boundary blurs, and the incident that used to take an hour of joining takes the minute it should. That is the goal of all of the above: not more logs, but logs that answer.