Promql

by grafana1ccacf29049fApache-2.0279 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. Covers `rate` vs `irate` vs `increase`, label matchers and regex, `sum / avg / topk / by / without` aggregation, classic + native `histogram_quantile`, ratios with divide-by-zero guards, `absent` / `changes` for staleness, time offsets and `predict_linear`, recording-rule naming, SLO + burn-rate math, and a cardinality-hunting playbook. Use when writing a metric query, fixing wrong p95s, building an error-budget alert, debugging "query is slow", finding the noisy label that blew up cardinality, or migrating a dashboard query to a recording rule — even when the user says "calculate the error rate", "p99 latency", "sum by service", "why is this query slow", or "what's filling Mimir" without naming PromQL.

AI-generated overview

Writes, validates and optimizes PromQL queries for Prometheus, Grafana Mimir and Grafana Cloud Metrics.

What it does
Provides PromQL patterns and workflows for writing and validating metric queries, including rate versus irate versus increase, label matchers, aggregations, histogram_quantile, ratio guards, staleness handling, offsets and SLO burn-rate math. It also covers converting slow dashboard expressions into recording rules and hunting high-cardinality labels. It produces query expressions, recording-rule YAML and validation steps against a Prometheus-compatible query API.
When to use it
Use when writing or fixing a metric query, correcting wrong p95 or p99 values, building an error-budget alert, debugging slow queries, finding labels that inflate cardinality, or migrating a dashboard query to a recording rule.
Requirements
A Prometheus, Grafana Mimir or Grafana Cloud Metrics query endpoint (with basic auth credentials for Grafana Cloud) and the bundled reference pattern library. Instructions only; no scripts are shipped.

PromQL Query Patterns

Docs: https://prometheus.io/docs/prometheus/latest/querying/basics/

PromQL returns either an instant vector, a range vector, or a scalar.

Golden rule: rate() / increase() require a range vector ≥ 4× the scrape interval. 60s scrape → use [5m] minimum.

Prerequisites

  • A Prometheus / Mimir / Grafana Cloud endpoint to query (/api/v1/query or via Grafana Explore)
  • The PromQL pattern library in references/patterns.md [blocked]

Common Workflows

1. Write + validate a query

bash
# 0. Point at your Prometheus/Mimir. For Grafana Cloud, use the metrics endpoint#    and add basic auth (-u "<metrics_user>:<token>") to each curl below.PROM=http://localhost:9090   # or https://prometheus-prod-XX.grafana.net/api/prom
# 1. Sketch the query — for "5xx error rate per service":EXPR='sum(rate(http_requests_total{status_code=~"5.."}[5m])) by (service)'
# 2. Validate syntax + that the metric/labels existcurl -sG --data-urlencode "query=${EXPR}" \  "$PROM/api/v1/query" | jq '.status, (.data.result|length)'# Expect: "success" and result count > 0. If 0 — check label spelling and scrape activity:curl -sG --data-urlencode "match[]=http_requests_total" "$PROM/api/v1/series" | jq '.data | length'
# 3. Sanity-check the magnitude — open Grafana Explore, paste the expr,#    confirm the values look right against a known ground truth (k6 run, log count, etc.)

2. Common patterns to copy

Per-status request rate (aggregate AFTER rate):

promql
sum(rate(http_requests_total{job="api"}[5m])) by (status_code)

p95 latency (must keep le in the inner aggregation):

promql
histogram_quantile(0.95,  sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))

Error rate with divide-by-zero guard:

promql
sum(rate(http_requests_total{status_code=~"5.."}[5m]))  / (sum(rate(http_requests_total[5m])) > 0)

Full library (recording rules, SLO burn-rate, offsets, cardinality hunt, native histograms): references/patterns.md [blocked].

3. Convert a slow dashboard query into a recording rule

yaml
# 1. Pick the slow expression, give it a recording-rule namegroups:  - name: http_request_rates    interval: 1m    rules:      - record: job:http_request_duration_p95:rate5m        expr: |          histogram_quantile(0.95,            sum(rate(http_request_duration_seconds_bucket[5m])) by (le, job))
bash
# 2. After rules load, verify the new metric existscurl -sG --data-urlencode "query=job:http_request_duration_p95:rate5m" \  "$PROM/api/v1/query" | jq '.data.result | length'   # → > 0
# 3. Verify it matches the original expression for at least one sample window# (Both queries should produce the same value at the same timestamp.)
# 4. Replace the dashboard panel expression with the recording-rule metric.

Common bugs

  • histogram_quantile returns NaN → forgot by (le) in the inner aggregation
  • "No data" → check the metric exists (/api/v1/series) and the window ≥ 4× scrape interval
  • Wrong rate magnitude → counter was aggregated before rate() (always rate() first)
  • Query timeout → series count too high; use topk(...) + a recording rule + drop high-cardinality labels (see references/patterns.md [blocked])

Resources

Source and attribution

Source:grafana/skillsinskills/grafana-core/promqlat commit1ccacf2

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from grafana/skills

React 19 Plugin Migration

grafana

Guides migration of a Grafana plugin to React 19 compatibility through ordered build, dependency and source-code steps.

Software Development279updated today

Plugin Bundle Size

grafana

Guides optimisation of Grafana app plugin bundle size using React.lazy, Suspense and webpack code splitting.

Software Development279updated today

Grafana Scenes

grafana

Builds Grafana plugin pages with the @grafana/scenes framework, covering scenes, panels, variables and drilldowns.

Software Development279updated today

Check Npm

grafana

Read-only audit of npm, yarn, or pnpm configuration for supply-chain hardening in a JS/TS repository.

Security279updated today

Mimir

grafana

Guides standing up and operating Grafana Mimir for scalable, multi-tenant, long-term Prometheus and OTLP metrics storage.

DevOps & Cloud279updated today

K6 Trend Analysis

grafana

Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for performance regression", "should I tighten thresholds", "are my tests degrading", "show me trends for test X", "analyze my k6 test runs", or "is my test getting slower". Also trigger when a user asks to check all tests in a project -- run this skill once per test and synthesize.

Awaiting classification279updated today