Prometheus Memory Keeps Growing After Adding a Label: Fixing Cardinality Explosion

Observability · Intermediate · 6 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Your Prometheus server started eating GBs of RAM and OOM-killing after one new metric label shipped. Here's how to find the offending series and cap cardinality for good.

The Problem

A service ships a one-line metric change on Tuesday. By Thursday your Prometheus memory keeps growing after adding that label: RSS goes from 3.2 GB steady-state to 15 GB, the pod gets OOM-killed, restarts, spends 9 minutes replaying the WAL, serves queries for 40 minutes, and OOMs again.

The change looked harmless:

httpDuration.WithLabelValues(route, method, status, tenantID).Observe(secs)

tenantID has 40,000 values. The histogram has 12 buckets plus _sum and _count. 40,000 × 12 routes × 14 = 6.7 million new time series. At roughly 1.5–3 KB of head memory per active series (labels, symbol table entries, open chunk, postings index), that's 10–20 GB — all of it resident, all of it before you store a single extra day of data.

Nothing errors. Ingestion succeeds. Prometheus just quietly grows until the kernel kills it.

Why the Obvious Fix Falls Short

The first three reactions are all reasonable and all wrong:

"Cut retention." Retention controls compacted blocks on disk. Head memory holds currently active series — the last ~2 hours of ingestion. Dropping retention from 15 days to 2 frees disk, not RAM. Your OOM is unchanged.

"Scrape less often." Going from 15s to 60s cuts samples per series by 4×, but samples are the cheap part: a sample is ~1.7 bytes after XOR compression inside an already-allocated chunk. The expensive part is the per-series overhead — the label set, its entry in the inverted index, and the 128-sample chunk buffer that exists whether you write to it 4 times a minute or once. You save maybe 10%, and now your alerts are 4× slower to fire.

"Give it more memory." This buys time proportional to a number you don't control. Tenant count grows; so does the memory. Worse, larger heads make WAL replay slower, so each OOM restart takes longer, and if replay itself needs more memory than the limit, you're in an unrecoverable crash loop where Prometheus never becomes ready.

The insight everyone misses: Prometheus memory is a function of unique label combinations, not of traffic volume or time range. A metric scraped once with 5M distinct label sets costs more than a metric scraped every second with 50.

How It Actually Works

Every unique combination of metric name + label values is a distinct time series with its own ID, its own entry in the in-memory inverted index, and its own open chunk. Adding a label doesn't add a column — it multiplies the series count by that label's cardinality.

flowchart TD
    A["Scrape: http_duration_bucket{route,method,status,tenant,le}"] --> B["Hash the full label set"]
    B --> C{"Series ID exists\nin head index?"}
    C -->|Yes| D["Append 1.7 bytes\nto existing open chunk"]
    C -->|No| E["Allocate new series:\nlabel strings + postings\nentry + 128-sample chunk"]
    E --> F["Head series count +1\n~1.5-3 KB RAM"]
    F --> G["Inverted index:\ntenant=abc -> [series IDs]\n40k new postings lists"]
    D --> H["Head block\n(last ~2h, all in RAM)"]
    F --> H
    H -->|"every 2h"| I["Compact to disk block"]
    H -.->|"series with no new samples\nlinger up to ~3h"| J["Memory not freed immediately"]

Two consequences fall straight out of this diagram:

  1. A histogram multiplies cardinality by bucket count. Adding a 40k-value label to a counter costs 40k series; adding it to a 12-bucket histogram costs ~560k.
  2. Churn is as bad as cardinality. If a label value changes on every deploy (pod=api-7d9f-x2k4, commit_sha=..., job_id=<uuid>), each rollout creates a fresh generation of series while the old ones sit in the head for hours. Steady-state series count looks fine; memory sawtooths upward.

Diagnose it with these three queries:

# Total active series — the number that predicts your RSS
prometheus_tsdb_head_series

# Are you creating series faster than they age out?
rate(prometheus_tsdb_head_series_created_total[10m])

# Which metric is responsible?
topk(10, count by (__name__)({__name__=~".+"}))

# Which label inside that metric is unbounded?
count(count by (tenant_id) (http_request_duration_seconds_bucket))

Prometheus 2.14+ also ships /tsdb-status in the web UI, which lists top label-value counts and top series-count-by-metric directly.

Before and After

Before — unbounded identifiers baked into labels:

// BAD: tenantID (40k values) x 12 buckets = ~500k series per route
var httpDuration = prometheus.NewHistogramVec(
    prometheus.HistogramOpts{
        Name:    "http_request_duration_seconds",
        Buckets: prometheus.DefBuckets,
    },
    []string{"route", "method", "status", "tenant_id"},
)

func handle(w http.ResponseWriter, r *http.Request) {
    start := time.Now()
    next.ServeHTTP(w, r)
    // r.URL.Path is the RAW path: /orders/8831, /orders/8832, ...
    // Second unbounded dimension hiding in plain sight.
    httpDuration.WithLabelValues(
        r.URL.Path, r.Method, status, tenantID(r),
    ).Observe(time.Since(start).Seconds())
}

After — every label has a small, enumerable domain; the high-cardinality identity moves to exemplars and logs:

// FIXED: plan_tier has 4 values, route is the TEMPLATE not the path.
// Series count: 12 routes x 5 statuses x 4 tiers x 14 = ~3,400. Bounded forever.
var httpDuration = prometheus.NewHistogramVec(
    prometheus.HistogramOpts{
        Name:    "http_request_duration_seconds",
        Buckets: prometheus.DefBuckets,
    },
    []string{"route", "method", "status", "plan_tier"},
)

func handle(w http.ResponseWriter, r *http.Request) {
    start := time.Now()
    next.ServeHTTP(w, r)
    obs := httpDuration.WithLabelValues(
        chi.RouteContext(r.Context()).RoutePattern(), // "/orders/{id}"
        r.Method, status, planTier(r),
    )
    // Per-request identity travels as an exemplar (stored separately,
    // fixed-size circular buffer) — not as a new time series.
    if e, ok := obs.(prometheus.ExemplarObserver); ok {
        e.ObserveWithExemplar(time.Since(start).Seconds(),
            prometheus.Labels{"trace_id": traceID(r), "tenant_id": tenantID(r)})
    } else {
        obs.Observe(time.Since(start).Seconds())
    }
}

And a server-side guardrail so one bad deploy can never OOM you again:

scrape_configs:
  - job_name: api
    sample_limit: 5000          # reject the whole scrape if a target explodes
    label_limit: 12
    label_value_length_limit: 128
    metric_relabel_configs:
      - source_labels: [__name__, tenant_id]
        regex: 'http_request_duration_seconds_.*;.+'
        action: drop            # emergency valve, no app redeploy needed

When NOT to Use This

Gotchas

Key takeaway: Prometheus memory scales with the number of unique label combinations currently in the head block, not with scrape interval or retention — so only removing or bounding label values fixes an OOM.

Real-world challenge

After a deploy, your Prometheus instance's RSS climbs from 4 GB to 14 GB over two hours and then OOMs; it restarts, replays the WAL for 8 minutes, and repeats. Nothing in the app changed except a new metric for background jobs. Grafana dashboards for older metrics still work between crashes. How do you confirm the cause and stop the crash loop without redeploying every service?

1. Confirm it's series count, not query load.

prometheus_tsdb_head_series
rate(prometheus_tsdb_head_series_created_total[5m])

A steep creation rate with flat sample ingestion means new label combinations, not more traffic.

2. Find the offender.

topk(10, count by (__name__)({__name__=~".+"}))

Then for the top metric, find which label is unbounded:

count(count by (job_id) (background_job_duration_seconds_bucket))

If that returns ~400k, job_id is a UUID.

3. Stop the bleeding at the scrape config — no app redeploy needed:

metric_relabel_configs:
  - source_labels: [__name__]
    regex: 'background_job_duration_seconds_.*'
    action: drop

Or keep the metric and drop just the label with action: labeldrop, regex: job_id (note this may collide series — for histograms, dropping the metric entirely is safer until the app is fixed).

4. Recover. The old series stay in the head for up to ~3 hours (two block periods) after their last sample, so memory won't drop instantly. If it's still OOMing, raise the memory limit temporarily so the WAL replay can finish, or delete the WAL and accept the data loss.

5. Fix properly in the app: move job_id to a log line or a trace exemplar, keep only job_type (bounded) as a label.