Prometheus Memory Keeps Growing After Adding a Label: Fixing Cardinality Explosion
Observability · Intermediate · 6 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Your Prometheus server started eating GBs of RAM and OOM-killing after one new metric label shipped. Here's how to find the offending series and cap cardinality for good.
The Problem
A service ships a one-line metric change on Tuesday. By Thursday your Prometheus memory keeps growing after adding that label: RSS goes from 3.2 GB steady-state to 15 GB, the pod gets OOM-killed, restarts, spends 9 minutes replaying the WAL, serves queries for 40 minutes, and OOMs again.
The change looked harmless:
httpDuration.WithLabelValues(route, method, status, tenantID).Observe(secs)
tenantID has 40,000 values. The histogram has 12 buckets plus _sum and _count. 40,000 × 12 routes × 14 = 6.7 million new time series. At roughly 1.5–3 KB of head memory per active series (labels, symbol table entries, open chunk, postings index), that's 10–20 GB — all of it resident, all of it before you store a single extra day of data.
Nothing errors. Ingestion succeeds. Prometheus just quietly grows until the kernel kills it.
Why the Obvious Fix Falls Short
The first three reactions are all reasonable and all wrong:
"Cut retention." Retention controls compacted blocks on disk. Head memory holds currently active series — the last ~2 hours of ingestion. Dropping retention from 15 days to 2 frees disk, not RAM. Your OOM is unchanged.
"Scrape less often." Going from 15s to 60s cuts samples per series by 4×, but samples are the cheap part: a sample is ~1.7 bytes after XOR compression inside an already-allocated chunk. The expensive part is the per-series overhead — the label set, its entry in the inverted index, and the 128-sample chunk buffer that exists whether you write to it 4 times a minute or once. You save maybe 10%, and now your alerts are 4× slower to fire.
"Give it more memory." This buys time proportional to a number you don't control. Tenant count grows; so does the memory. Worse, larger heads make WAL replay slower, so each OOM restart takes longer, and if replay itself needs more memory than the limit, you're in an unrecoverable crash loop where Prometheus never becomes ready.
The insight everyone misses: Prometheus memory is a function of unique label combinations, not of traffic volume or time range. A metric scraped once with 5M distinct label sets costs more than a metric scraped every second with 50.
How It Actually Works
Every unique combination of metric name + label values is a distinct time series with its own ID, its own entry in the in-memory inverted index, and its own open chunk. Adding a label doesn't add a column — it multiplies the series count by that label's cardinality.
flowchart TD
A["Scrape: http_duration_bucket{route,method,status,tenant,le}"] --> B["Hash the full label set"]
B --> C{"Series ID exists\nin head index?"}
C -->|Yes| D["Append 1.7 bytes\nto existing open chunk"]
C -->|No| E["Allocate new series:\nlabel strings + postings\nentry + 128-sample chunk"]
E --> F["Head series count +1\n~1.5-3 KB RAM"]
F --> G["Inverted index:\ntenant=abc -> [series IDs]\n40k new postings lists"]
D --> H["Head block\n(last ~2h, all in RAM)"]
F --> H
H -->|"every 2h"| I["Compact to disk block"]
H -.->|"series with no new samples\nlinger up to ~3h"| J["Memory not freed immediately"]
Two consequences fall straight out of this diagram:
- A histogram multiplies cardinality by bucket count. Adding a 40k-value label to a counter costs 40k series; adding it to a 12-bucket histogram costs ~560k.
- Churn is as bad as cardinality. If a label value changes on every deploy (
pod=api-7d9f-x2k4,commit_sha=...,job_id=<uuid>), each rollout creates a fresh generation of series while the old ones sit in the head for hours. Steady-state series count looks fine; memory sawtooths upward.
Diagnose it with these three queries:
# Total active series — the number that predicts your RSS
prometheus_tsdb_head_series
# Are you creating series faster than they age out?
rate(prometheus_tsdb_head_series_created_total[10m])
# Which metric is responsible?
topk(10, count by (__name__)({__name__=~".+"}))
# Which label inside that metric is unbounded?
count(count by (tenant_id) (http_request_duration_seconds_bucket))
Prometheus 2.14+ also ships /tsdb-status in the web UI, which lists top label-value counts and top series-count-by-metric directly.
Before and After
Before — unbounded identifiers baked into labels:
// BAD: tenantID (40k values) x 12 buckets = ~500k series per route
var httpDuration = prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Buckets: prometheus.DefBuckets,
},
[]string{"route", "method", "status", "tenant_id"},
)
func handle(w http.ResponseWriter, r *http.Request) {
start := time.Now()
next.ServeHTTP(w, r)
// r.URL.Path is the RAW path: /orders/8831, /orders/8832, ...
// Second unbounded dimension hiding in plain sight.
httpDuration.WithLabelValues(
r.URL.Path, r.Method, status, tenantID(r),
).Observe(time.Since(start).Seconds())
}
After — every label has a small, enumerable domain; the high-cardinality identity moves to exemplars and logs:
// FIXED: plan_tier has 4 values, route is the TEMPLATE not the path.
// Series count: 12 routes x 5 statuses x 4 tiers x 14 = ~3,400. Bounded forever.
var httpDuration = prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Buckets: prometheus.DefBuckets,
},
[]string{"route", "method", "status", "plan_tier"},
)
func handle(w http.ResponseWriter, r *http.Request) {
start := time.Now()
next.ServeHTTP(w, r)
obs := httpDuration.WithLabelValues(
chi.RouteContext(r.Context()).RoutePattern(), // "/orders/{id}"
r.Method, status, planTier(r),
)
// Per-request identity travels as an exemplar (stored separately,
// fixed-size circular buffer) — not as a new time series.
if e, ok := obs.(prometheus.ExemplarObserver); ok {
e.ObserveWithExemplar(time.Since(start).Seconds(),
prometheus.Labels{"trace_id": traceID(r), "tenant_id": tenantID(r)})
} else {
obs.Observe(time.Since(start).Seconds())
}
}
And a server-side guardrail so one bad deploy can never OOM you again:
scrape_configs:
- job_name: api
sample_limit: 5000 # reject the whole scrape if a target explodes
label_limit: 12
label_value_length_limit: 128
metric_relabel_configs:
- source_labels: [__name__, tenant_id]
regex: 'http_request_duration_seconds_.*;.+'
action: drop # emergency valve, no app redeploy needed
When NOT to Use This
- You genuinely need per-entity data. If the question is "what was p99 latency for customer X last Tuesday," Prometheus is the wrong store. Use traces (Tempo/Jaeger) with exemplar links, or a columnar log store (ClickHouse, Loki with structured metadata). Don't try to make a TSDB behave like an event store.
- Cardinality is moderate but you need it. 5k–50k extra series is fine — that's tens of MB. Only start cutting when a single metric family is producing hundreds of thousands of series. Blanket "no labels ever" rules cause teams to ship metrics that answer nothing.
- You've already outgrown a single server. If your legitimate, bounded cardinality is 20M+ series, the fix isn't relabelling — it's Thanos/Mimir/VictoriaMetrics with horizontal sharding, or Prometheus agent mode with remote-write.
- Short-lived spikes.
sample_limitwill hard-drop an entire target's scrape, blinding you exactly when something is going wrong. Prefer targetedmetric_relabel_configsdrops over a global sample limit on critical jobs.
Gotchas
- Memory doesn't drop when you deploy the fix. Dead series stay in the head until the block containing them is compacted out — typically 2–3 hours. Don't conclude the fix failed at minute 10.
- WAL replay is the real killer. A restart must load the entire head back into memory before serving. If your head needs 14 GB and your limit is 12 GB, Prometheus can never become ready. Temporarily raise the limit to get through replay, then apply relabelling.
labeldropon a histogram silently merges series. Two series that become identical after dropping a label both append to the same series with conflicting timestamps — you'll getout-of-order sampleerrors. Drop the whole metric first, fix the app, then re-enable.- Kubernetes labels churn by default.
pod,container_id,replicasetand anycommit_shalabel create a whole new series generation on every rollout. Deploying 30× a day with 200k series per generation is a cardinality problem even though no single label is unbounded. sum without (tenant_id)in a recording rule does not reduce storage. The raw series are still ingested and stored; you've only made the query cheaper. Cardinality must be cut at scrape time or before.- Error strings make terrible label values.
error="dial tcp 10.0.4.19:5432: connect: connection refused"embeds an IP and port — effectively unbounded. Map to a small enum (error_class="conn_refused") and log the detail.
Key takeaway: Prometheus memory scales with the number of unique label combinations currently in the head block, not with scrape interval or retention — so only removing or bounding label values fixes an OOM.
Real-world challenge
After a deploy, your Prometheus instance's RSS climbs from 4 GB to 14 GB over two hours and then OOMs; it restarts, replays the WAL for 8 minutes, and repeats. Nothing in the app changed except a new metric for background jobs. Grafana dashboards for older metrics still work between crashes. How do you confirm the cause and stop the crash loop without redeploying every service?
1. Confirm it's series count, not query load.
prometheus_tsdb_head_series
rate(prometheus_tsdb_head_series_created_total[5m])
A steep creation rate with flat sample ingestion means new label combinations, not more traffic.
2. Find the offender.
topk(10, count by (__name__)({__name__=~".+"}))
Then for the top metric, find which label is unbounded:
count(count by (job_id) (background_job_duration_seconds_bucket))
If that returns ~400k, job_id is a UUID.
3. Stop the bleeding at the scrape config — no app redeploy needed:
metric_relabel_configs:
- source_labels: [__name__]
regex: 'background_job_duration_seconds_.*'
action: drop
Or keep the metric and drop just the label with action: labeldrop, regex: job_id (note this may collide series — for histograms, dropping the metric entirely is safer until the app is fixed).
4. Recover. The old series stay in the head for up to ~3 hours (two block periods) after their last sample, so memory won't drop instantly. If it's still OOMing, raise the memory limit temporarily so the WAL replay can finish, or delete the WAL and accept the data loss.
5. Fix properly in the app: move job_id to a log line or a trace exemplar, keep only job_type (bounded) as a label.