← blog · September 4, 2026

Metrics, Logs, Traces, and Error Tracking: Why You Need All Four

Metrics, logs, traces, and error tracking do not measure the same thing. Each answers a different question at a different cost. A guide to when you need which, and the traps that come with each.

"Observability" gets talked about as if it were one thing, but underneath it are four distinct kinds of signal: metrics, logs, traces, and error tracking. Each answers a different question, at a different scale, and at a different cost. Trying to make one of them stand in for the others usually means looking at the wrong data during an actual incident.

Metrics: they show the trend, not the reason

A metric is a numeric series collected over time: how many requests per second, how many of them failed, how latency is distributed. The RED method (rate, errors, duration) and the USE method (utilization, saturation, errors) remain useful starting points precisely because they answer "which metric should I collect first" in one sentence.

The strength of a metric is that it is cheap. A counter or histogram reduces millions of requests to a single time series, so it can be stored for months and queried in milliseconds. That same strength is also its limit: a metric is an aggregate, it does not tell the story of one request. "Error rate jumped to three percent" is enough to fire an alert, but it says nothing about which user, which parameter, or which downstream call produced the failure.

The most common trap here is putting high-cardinality values into metric labels: user id, request id, full URL path. Time series databases like Prometheus store every label combination as a separate series, so putting a user id in a label means opening one series per user. Memory usage explodes and queries slow to a crawl. Anything with high cardinality belongs in a log or a trace, not in a metric label.

Logs: they give detail, at a real storage cost

A log is the record of a single event at a specific moment. It gives the detail a metric cannot: which user, which input, which exact error message. That detail comes with a cost in storage and search; a few log lines per request adds up to terabytes within days on a system with real traffic.

Unstructured, plain-text logs make this worse, because searching falls back to regular expressions and correlating across fields is close to impossible. Structured logs (JSON, key-value pairs) attach fields to every line, and the most important field is a request id or trace id. Without it, finding the handful of log lines that belong to the one request a user complained about is a needle-in-a-haystack problem, even with a good search tool.

Traces: following a request across services

When a request does not end in a single service, a gateway, a few microservices, a queue, a database, neither a metric nor a log alone can tell you where the total latency accumulated. A trace links every step (span) of a request together with timestamps and durations; the W3C Trace Context standard's traceparent header is the common format that carries this context between services.

The real value of a trace is finding the source of latency in a distributed system. "The request took 800 milliseconds" on its own is not actionable, but seeing in the trace that "650 of those milliseconds were spent in a synchronous call to a third service" gives you an actual target to optimize. The difficulty is that being traceable requires the code, or at least the libraries, to carry the context forward into every call. That context breaks easily when a message goes onto a queue or into a background job, because the message usually carries only the business payload, not the trace id. A broken context shows up as a gap in the middle of the trace, and the real source of latency stays invisible.

Tracing every single request is expensive at scale, so sampling is unavoidable. Head-based sampling decides at the start of a request, at random; it is simple to implement but likely to miss the rare slow request. Tail-based sampling decides after the whole trace has been collected, and always keeps traces that show an error or exceed a latency threshold; it is harder to run but it does not miss the very outliers you were looking for.

Error tracking: not a log, a grouping engine

Error tracking tools are often confused with logs, but their job is different. Sending every exception to a log collector as a separate line produces a thousand lines on a day when the same error repeats a thousand times, and it is left to a human to tell which line is new and which is a known, already-triaged issue. The real value of an error tracker is grouping, or fingerprinting: it collapses exceptions with the same stack trace into a single record and shows how many times it happened, which release introduced it, and which users it affected.

That grouping has a cost of its own. The fingerprinting algorithm sometimes merges two genuinely different errors under one record, for example two exceptions of the same type raised at different lines, and sometimes splits a minor variation of the same error, a changed variable name, a different third-party library version, into a separate record. That is why a new error tracking setup is worth testing against a handful of real examples rather than trusting the default grouping rules out of the box.

How the four work together in practice

In a real incident, the four signals come in, in order. A metric alert (error rate, p99 latency) is usually first to say that something changed. A trace narrows down which service or dependency is responsible. Logs from that service show exactly which input produced which error. The error tracker then merges that single event with others from the same root cause and answers whether this is a new problem or a known one already sitting in the backlog.

For that chain to work, one shared field is required: the trace id. The three signals other than metrics can be linked together as long as they carry the same trace id, on the log line, on the error record, on the trace's own spans. This is the actual reason OpenTelemetry gained traction: it collects metrics, logs, and traces separately, but through a single set of libraries that carry the same context across all three.

When not to invest in all four

On a small system running on a single server, a single service, with a few deploys a week at most, building a full observability stack, a metrics collector, a log aggregator, tracing infrastructure, and a separate error tracker, costs more in operational overhead than it returns. Structured logs plus a basic uptime check are usually enough there; even an error tracker is optional, because with one service it is not hard to tell where an error came from.

The investment starts paying off once there is more than one service and more than one deploy a day. From that point on, questions like "which deploy caused the error rate to go up" or "at which service boundary is the latency accumulating" cannot be answered by metrics or logs alone, and tracing becomes necessary. The order matters: structured logs first, then metrics and alerting, then error tracking, and tracing last. Setting up tracing before you have more than one service is buying complexity you do not need yet.