Prometheus Cardinality Troubleshooter

by grafana1ccacf29049fApache-2.0279 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid memory growth. Walks through tsdb status endpoints, per-metric and per-label drill-downs, common-culprit galleries, and remediation paths. Use when the user is *currently experiencing* a cardinality fire. For preventing cardinality issues at the source, route to prometheus-label-strategy. For post-ingest aggregation, route to adaptive-metrics. For DPM-specific analysis, route to dpm-finder.

Instructions onlyDevOps & Cloud

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
SKILL.md18.1 KBtext/markdown

Source and attribution

Source:grafana/skillsinskills/grafana-cloud/prometheus-cardinality-troubleshooterat commit1ccacf2

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from grafana/skills

React 19 Plugin Migration

grafana

Guides migration of a Grafana plugin to React 19 compatibility through ordered build, dependency and source-code steps.

Software Development279updated today

Plugin Bundle Size

grafana

Guides optimisation of Grafana app plugin bundle size using React.lazy, Suspense and webpack code splitting.

Software Development279updated today

Grafana Scenes

grafana

Builds Grafana plugin pages with the @grafana/scenes framework, covering scenes, panels, variables and drilldowns.

Software Development279updated today

Check Npm

grafana

Read-only audit of npm, yarn, or pnpm configuration for supply-chain hardening in a JS/TS repository.

Security279updated today

Mimir

grafana

Guides standing up and operating Grafana Mimir for scalable, multi-tenant, long-term Prometheus and OTLP metrics storage.

DevOps & Cloud279updated today

K6 Trend Analysis

grafana

Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for performance regression", "should I tighten thresholds", "are my tests degrading", "show me trends for test X", "analyze my k6 test runs", or "is my test getting slower". Also trigger when a user asks to check all tests in a project -- run this skill once per test and synthesize.

Awaiting classification279updated today