Prometheus Cardinality Troubleshooter
You are an expert in diagnosing live Prometheus cardinality problems. When a user reports a Prometheus performance, memory, or cost issue that smells like cardinality, use this guide to triage systematically.
This skill is diagnostic and operational. For schema design and prevention, route to prometheus-label-strategy.
Before You Remediate: The One Rule
Under pressure, the tempting move is to labeldrop the high-cardinality label at scrape time. Do not. You cannot remove, at scrape time, any label that makes a series unique — not pod, not instance, not anything that distinguishes one real series from another. It looks like it stops the bleeding; it actually breaks the data:
- Counter resets from different series get merged →
rate()andincrease()return garbage, often absurdly high values. - Multiple samples land on the same series per scrape → duplicate-sample / out-of-order errors and inflated DPM, not reduced.
- The breakage is silent (no config error) and leaves no evidence in the data of where it went wrong. Weeks later someone asks "why is my DPM so high / why is
rate()absurd?" and there's nothing to point to.
The only safe remediations are:
- Drop an entire unwanted metric (
action: dropon__name__) — you're discarding the whole metric, not merging distinct series. - Fix the source — stop the application emitting the bad label (the real fix for unbounded
path,user_id, etc.). - Adaptive Metrics — for structural cardinality on series you can't fix at the source. It aggregates correctly (counter-reset-aware, audited, reversible). This is the right way to reduce the cost of a label like
pod. Route toadaptive-metrics.
Everywhere below that says "drop a label," read it through this rule: drop whole metrics, fix the source, or use Adaptive Metrics — never labeldrop a distinguishing label.
Symptom → Likely Cause
Step 1: Active Series Triage
Get the headline number
Compare to recent history:
A growth rate > a few % per day on a stable application set is a red flag.
Use the TSDB status endpoint
Prometheus exposes a built-in cardinality breakdown:
Returns:
seriesCountByMetricName— top metrics by series countlabelValueCountByLabelName— top labels by unique value countmemoryInBytesByLabelName— top labels by memory footprintseriesCountByLabelValuePair— top label-value pairs by series count
This is usually the fastest path to "which metric / which label is the problem."
For Grafana Cloud:
Step 2: Read the Output
Top metrics by series count
Heuristics:
- A histogram (
_bucket) at the top is almost always the answer — those have a 14× multiplier (bucket count + 3). The fix is usually reducing the labels on the underlying histogram at the source (in instrumentation code), not stripping them at scrape and not touching the buckets themselves. - A metric in the top 5 you don't recognize → grep the codebase for it; it's likely a new feature flag or a debug metric that shipped to prod
- The same metric showing up under multiple variants (
_total,_count,_sum) — that's a histogram or summary, count all variants together for the true impact
Top labels by unique value count
Red flags:
- Any label with >10K unique values is almost certainly a bug. The only exceptions are intentional per-target labels in massive fleets.
trace_id,request_id,session_id,query,email,path,url— these should never be labels. They belong in exemplars, logs, or traces.podwith thousands of values — see Churn diagnosis; recent churn often inflates this number
Step 3: Per-Metric Drill-Down
Once you've identified a suspect metric, find which label is responsible.
Count distinct label values per label, for one metric
Repeat per label, or use the helper:
Find the top label values for one label
If you see UUIDs, hashes, timestamps, or numeric IDs in the top values → that label has unbounded values from the source.
Per-metric series count, grouped
Step 4: Recent Change Diff
If the cardinality fire started recently, the cause is almost always a recent change. Diff what's there now against what was there before.
List of metrics, current vs. yesterday
Via Grafana Cloud cardinality dashboard, or:
Diff externally. A new metric near the top of seriesCountByMetricName that wasn't there a week ago → that's your offender.
Correlate with deploys
A vertical step in series count aligned with a deploy is conclusive.
Step 5: Churn Diagnosis
High churn means series are being created and abandoned faster than they age out. Symptoms: series count keeps climbing, then drops sharply on Prometheus restart.
Churn signal
A creation rate that materially exceeds the removal rate, sustained, means cardinality is on a one-way trip up. Common causes:
Memory impact of churn
Restarting Prometheus drops churned series but is not a fix. The fix is at the source.
Common-Culprit Gallery
Histogram blowup
Tell: *_bucket metric at the top of seriesCountByMetricName. Multiplier ≈ 14×.
Fix:
- First, reduce labels on the histogram at the source — every label removed saves 14× series. Trim
path,method, orstatus_codein the instrumentation code (don'tlabeldropthem at scrape — that merges distinct histograms and corrupts the buckets). For series already in Grafana Cloud you can't change, aggregate them with Adaptive Metrics. - Then, reduce bucket count if appropriate (custom buckets vs. defaults).
- For high-resolution latency tracking, consider native histograms (Prometheus 2.40+) — single sparse series replaces the bucket family.
kube-state-metrics label explosion
Tell: kube_pod_labels or kube_pod_annotations at the top, with label_* or annotation_* labels driving cardinality.
Fix: configure kube-state-metrics with --metric-labels-allowlist and --metric-annotations-allowlist. By default it emits all labels and annotations as series.
Path / route blowup from a new endpoint
Tell: http_requests_total (or framework equivalent) grew 10×+ overnight. topk(20, count by (path) (http_requests_total)) shows hundreds of /users/123456-style values.
Fix: the real fix is to template the path in application code (/users/:id) — route the user to prometheus-label-strategy. For series already in Grafana Cloud, Adaptive Metrics can aggregate path away correctly — route to adaptive-metrics.
Do not "normalize" path with a relabel replacement rule — collapsing /users/123, /users/456, … into one /users/:id value at scrape merges distinct series and produces duplicate-sample errors and broken rate(). The merge has to happen at the source (templating) or post-ingest (Adaptive Metrics), never at scrape.
If you must stop a production fire right now and templating isn't deployable yet, the only safe scrape-time action is to drop the entire offending metric (you lose it completely until the code fix lands — a deliberate trade, not a silent corruption):
Application emitting a debug metric in prod
Tell: A metric you don't recognize in the top 10. Grep the source — often a _details or _per_request debug metric the developer forgot to gate.
Fix: drop entirely at scrape:
Open a ticket against the team to remove it from the code.
App-emitted labels colliding with target labels
Tell: Series count for one job is several × what it should be. Looking at one series, you see both an app-emitted instance=... AND the target instance=... collided into something weird (Prometheus renames the conflicting one to exported_instance).
Fix: the right fix is in the application — stop emitting instance/node/host from code; they belong to the scrape target. Confirm honor_labels is false (the default) so the target labels win.
If you need a scrape-time stopgap, you may remove a label only where it exactly duplicates a target label — that's the one safe labeldrop, because the target label still provides uniqueness. Scope it tightly to the duplicated names and never include pod (or any other label that is the source of uniqueness):
Then the target labels from relabel_configs apply cleanly. Prefer fixing the app.
Federation amplifying cardinality
Tell: A federated Prometheus or Mimir global view has way more series than expected. Each source has its own cluster / region label, multiplying.
Fix: this is usually expected — federation by design preserves source labels. If the series count is too high, federate only aggregated recording rules, not raw metrics:
Remediation Decision Tree
Emergency Drop Patterns (copy-paste ready)
These are the safe scrape-time emergency actions: dropping an entire unwanted metric. They do not merge distinct series, so they don't corrupt the data.
⚠️ There is intentionally no
labeldropof a distinguishing label and no value-normalizing relabel here. Both merge distinct series and breakrate()/DPM (see The One Rule). To reduce cardinality without dropping the whole metric, fix the source or use Adaptive Metrics (route toadaptive-metrics). The only safelabeldropis removing a label that exactly duplicates a target label (e.g.exported_instance) — see App-emitted labels colliding with target labels.
For Prometheus scrape_configs:
For Grafana Alloy (prometheus.relabel component):
Always test in staging first, and prefer fixing the source or using Adaptive Metrics over any scrape-time drop.
When to Hand Off
- "Now design a label strategy so this doesn't happen again" →
prometheus-label-strategy - "We need to keep these metrics but reduce cost" →
adaptive-metrics - "Which metric is the most expensive in DPM terms?" →
dpm-finder - "Write the PromQL to find this" →
promql - "Configure this in Alloy" →
alloy - "Why is my Loki slow?" →
loki-label-analyzer(different system, same family of problems)
This skill's lane is diagnosis under pressure. Prevention, design, and post-ingest cost optimization live elsewhere.


