← blog · October 9, 2026

Cardinality Explosions in Prometheus: Choosing Labels, What a Series Costs and Stopping It at the Source

Every unique label combination is its own series, and one unbounded label can run Prometheus out of memory. Measuring series counts, the closed-set rule for labels, scrape limits, labeldrop collisions, what histograms cost and what recording rules do not fix.

In Prometheus, every unique combination of a metric name and its label values is a separate time series. http_requests_total{method="GET", route="/api/orders", status="200"} is one series, and the same metric with status="500" is another. The model is cheap and fast until someone puts an unbounded value into a label: a user ID, a raw URL path, an error message. From then on the series count grows with traffic, Prometheus memory grows along with it, queries slow down, and one night the server restarts out of memory. This post covers where series come from, how to measure them, and where to stop them.

How series multiply

The series count is the product of the number of values each label takes. An application with 5 methods, 40 routes and 8 status codes produces at most 1,600 series; scraped across 20 pods it becomes 32,000, because the pod or instance label is a multiplier too. That is still fine. Replace the route with the raw path (/api/orders/81723) and one of the factors becomes infinite.

Histograms make the multiplier bigger. For each label combination, a histogram produces one _bucket series per bucket, plus _sum and _count. The Go client's 11 default boundaries become 12 buckets once +Inf is added, so each combination is 14 series. Had the example above been a histogram instead of a counter, it would be 448,000 series rather than 32,000.

The third source is churn. Every deployment brings pods with new names; the old pods' series stop receiving samples, but they do not disappear right away either. Prometheus keeps the last few hours of data in its in-memory head block, and stale series live there until the next compaction. A service that deploys several times a day, or a job that starts a new pod on every run, multiplies the active series count even at constant traffic.

Measure first

Prometheus reports its own total of active series: prometheus_tsdb_head_series. Which metric and which label are to blame is listed on the Status > TSDB Status page of the UI (or the /api/v1/status/tsdb endpoint): metric names with the most series, label names with the most values, labels holding the most memory, and the label value pairs behind the most series. For data on disk, promtool tsdb analyze runs the same analysis in more detail and shows churn as well.

On the query side:

topk(10, count by (__name__) ({__name__=~".+"}))
count(count by (path) (http_requests_total))
sum by (job) (scrape_series_added)

The first returns the metrics with the most series, but it touches every series, so do not run it often. The second counts how many distinct values a label takes. The third shows, per job, how many new series each scrape added; if it stays high at constant traffic, a churning label is hiding somewhere.

Choosing labels: the closed-set rule

A label value must come from a small, closed set that is known in advance. Method, status class, route template, region and queue name are fine. User IDs, tenant IDs that can run into the thousands, email addresses, request and session IDs, raw URL paths, query strings, error message text, timestamps and IP addresses are not.

For routes, use the template your framework matched (/api/orders/{id}), not the raw path. Collect unmatched requests under one fixed value (route="unmatched"): otherwise a scanner bot trying random paths opens a new series for every attempt and can take your monitoring down simply by sending traffic. For status codes, the 2xx, 4xx, 5xx class is enough for most dashboards.

Per-event detail (which user, which request, which error text) belongs in logs or traces, not in metrics. Exemplars can link the two: a trace ID is attached to a histogram observation, and from the dashboard you jump straight to that request's trace. Depending on your version, exemplar storage in Prometheus may be a feature you have to enable separately.

Stop it at the source: scrape limits

So that the server does not fall over before you notice a bad label in the application, put upper bounds into the scrape configuration:

scrape_configs:
  - job_name: app
    sample_limit: 20000
    label_limit: 30
    label_value_length_limit: 200
    metric_relabel_configs:
      - source_labels: [__name__]
        regex: "http_client_request_duration_seconds_bucket"
        action: drop
      - regex: "request_id|session_id"
        action: labeldrop

When sample_limit is exceeded, the whole scrape counts as failed: none of the target's metrics arrive and up drops to zero. That is a deliberate choice, a loud failure instead of partial data. But it only helps if you alert on up == 0; watch the prometheus_target_scrapes_exceeded_sample_limit_total counter as well, so the reason is obvious right away.

metric_relabel_configs runs after the scrape and before storage; a dropped series is never written. The first rule above discards the buckets of a histogram nobody looks at, while _sum and _count remain, so you can still compute an average. Be careful when dropping labels: two series that differ only by that label collapse into the same identity after labeldrop, and Prometheus rejects the second sample with the same timestamp. One of the samples is lost, and the only places you see it are the prometheus_target_scrapes_sample_duplicate_timestamp_total counter and the log. If you are not sure the label you drop does not distinguish any series, fix it in the application instead.

In Kubernetes, ready-made exporters such as the kubelet and cAdvisor emit hundreds of metrics your dashboards never use. Listing the metrics that appear in dashboards and alert rules and dropping the rest at scrape time cuts the active series count noticeably in most clusters.

What recording rules fix, and what they do not

Recording rules precompute an expensive aggregation:

groups:
  - name: http
    rules:
      - record: job_route:http_requests:rate5m
        expr: sum by (job, route, code) (rate(http_requests_total[5m]))

Dashboards get faster, but the raw series are still stored, so the memory problem is not solved; new series are added on top. The ways to actually reduce the series count are removing the label in the application, dropping it at scrape time, or a separate layer that aggregates before writing to remote storage.

Think about histograms separately

Histograms are where cardinality grows fastest. Instead of using the default bucket list blindly, choose buckets around your SLO thresholds: with a 300 ms target, dense buckets around that threshold and sparse ones further out do the job. Give histograms fewer labels than counters. Newer Prometheus versions offer native histograms, which keep a single series instead of one per bucket; before relying on them, check whether they are stable in your version and supported by your client library.

Pitfalls

The bill on remote storage. Managed metrics services usually charge by active series. A label that is a memory problem locally turns straight into an invoice remotely, and by the time anyone notices, half the month may be gone.

Short-lived jobs. Scheduled jobs that start a new pod on every run are only scraped a few times, but each one leaves a fresh set of series behind. Keep metrics from such jobs without the pod label at the point where they are collected.

Library defaults. Some HTTP middleware labels requests by raw path instead of route, or puts the target URL into a label on the client side. When you add a new library, make its /metrics output the first thing you look at.

Values that do not look unbounded. Status codes look like a closed set, but a proxy that passes through whatever the backend returns can produce hundreds of non-standard codes. If queues are created per tenant, queue names are unbounded too.

Where Prometheus is the wrong tool

Questions such as "how many requests did each user send" or "which customer's endpoint is slow" are high-cardinality data by nature. Instead of squeezing them into Prometheus labels, write them to structured logs or to a columnar store built for event data. Let metrics tell you that something is wrong and how wrong it is; let logs and traces tell you who it is wrong for.

A short checklist

Every label has a small, closed set of values. Routes are templates, not raw paths, and unmatched requests are collected under a single value. Histogram buckets and labels are chosen on purpose. Every scrape job has a sample_limit and label limits, and exceeding them raises an alert. prometheus_tsdb_head_series sits on a dashboard and alerts on unexpected growth. When a new library or exporter is added, the first thing checked is how many series it produces. Once this is in place, cardinality stops being an out-of-memory event at midnight and becomes an ordinary topic caught in a pull request review.