Promql

作者 grafana1ccacf29049fApache-2.0279 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. Covers `rate` vs `irate` vs `increase`, label matchers and regex, `sum / avg / topk / by / without` aggregation, classic + native `histogram_quantile`, ratios with divide-by-zero guards, `absent` / `changes` for staleness, time offsets and `predict_linear`, recording-rule naming, SLO + burn-rate math, and a cardinality-hunting playbook. Use when writing a metric query, fixing wrong p95s, building an error-budget alert, debugging "query is slow", finding the noisy label that blew up cardinality, or migrating a dashboard query to a recording rule — even when the user says "calculate the error rate", "p99 latency", "sum by service", "why is this query slow", or "what's filling Mimir" without naming PromQL.

AI 產生的概覽

撰寫、驗證並最佳化用於 Prometheus、Grafana Mimir 與 Grafana Cloud Metrics 的 PromQL 查詢。

功能
提供 PromQL 模式與工作流程,用於撰寫與驗證指標查詢,涵蓋 rate、irate 與 increase 的差異、標籤比對器、聚合、histogram_quantile、除零保護、過期資料處理、時間位移與 SLO 燃燒率計算。也涵蓋將緩慢的儀表板運算式轉換為記錄規則,以及追查高基數標籤。產出查詢運算式、記錄規則 YAML,以及針對相容於 Prometheus 的查詢 API 的驗證步驟。
適用情境
適用於撰寫或修正指標查詢、修正錯誤的 p95 或 p99 數值、建立錯誤預算告警、偵錯緩慢查詢、找出導致基數暴增的標籤,或將儀表板查詢遷移為記錄規則。
執行需求
需要 Prometheus、Grafana Mimir 或 Grafana Cloud Metrics 查詢端點(Grafana Cloud 需基本驗證憑證)以及隨附的參考模式庫。僅為說明文件,不附帶指令碼。

PromQL Query Patterns

Docs: https://prometheus.io/docs/prometheus/latest/querying/basics/

PromQL returns either an instant vector, a range vector, or a scalar.

Golden rule: rate() / increase() require a range vector ≥ 4× the scrape interval. 60s scrape → use [5m] minimum.

Prerequisites

  • A Prometheus / Mimir / Grafana Cloud endpoint to query (/api/v1/query or via Grafana Explore)
  • The PromQL pattern library in references/patterns.md [blocked]

Common Workflows

1. Write + validate a query

bash
# 0. Point at your Prometheus/Mimir. For Grafana Cloud, use the metrics endpoint#    and add basic auth (-u "<metrics_user>:<token>") to each curl below.PROM=http://localhost:9090   # or https://prometheus-prod-XX.grafana.net/api/prom
# 1. Sketch the query — for "5xx error rate per service":EXPR='sum(rate(http_requests_total{status_code=~"5.."}[5m])) by (service)'
# 2. Validate syntax + that the metric/labels existcurl -sG --data-urlencode "query=${EXPR}" \  "$PROM/api/v1/query" | jq '.status, (.data.result|length)'# Expect: "success" and result count > 0. If 0 — check label spelling and scrape activity:curl -sG --data-urlencode "match[]=http_requests_total" "$PROM/api/v1/series" | jq '.data | length'
# 3. Sanity-check the magnitude — open Grafana Explore, paste the expr,#    confirm the values look right against a known ground truth (k6 run, log count, etc.)

2. Common patterns to copy

Per-status request rate (aggregate AFTER rate):

promql
sum(rate(http_requests_total{job="api"}[5m])) by (status_code)

p95 latency (must keep le in the inner aggregation):

promql
histogram_quantile(0.95,  sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))

Error rate with divide-by-zero guard:

promql
sum(rate(http_requests_total{status_code=~"5.."}[5m]))  / (sum(rate(http_requests_total[5m])) > 0)

Full library (recording rules, SLO burn-rate, offsets, cardinality hunt, native histograms): references/patterns.md [blocked].

3. Convert a slow dashboard query into a recording rule

yaml
# 1. Pick the slow expression, give it a recording-rule namegroups:  - name: http_request_rates    interval: 1m    rules:      - record: job:http_request_duration_p95:rate5m        expr: |          histogram_quantile(0.95,            sum(rate(http_request_duration_seconds_bucket[5m])) by (le, job))
bash
# 2. After rules load, verify the new metric existscurl -sG --data-urlencode "query=job:http_request_duration_p95:rate5m" \  "$PROM/api/v1/query" | jq '.data.result | length'   # → > 0
# 3. Verify it matches the original expression for at least one sample window# (Both queries should produce the same value at the same timestamp.)
# 4. Replace the dashboard panel expression with the recording-rule metric.

Common bugs

  • histogram_quantile returns NaN → forgot by (le) in the inner aggregation
  • "No data" → check the metric exists (/api/v1/series) and the window ≥ 4× scrape interval
  • Wrong rate magnitude → counter was aggregated before rate() (always rate() first)
  • Query timeout → series count too high; use topk(...) + a recording rule + drop high-cardinality labels (see references/patterns.md [blocked])

Resources

來源與署名

來源:grafana/skills位於skills/grafana-core/promql提交1ccacf2

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 grafana/skills 的技能

React 19 Plugin Migration

grafana

指導將 Grafana 外掛遷移至 React 19 相容,依序完成建置、相依性與原始碼修改步驟。

Software Development279今天更新

Plugin Bundle Size

grafana

指導使用 React.lazy、Suspense 與 webpack 程式碼分割來最佳化 Grafana 應用程式外掛的打包體積。

Software Development279今天更新

Grafana Scenes

grafana

使用 @grafana/scenes 框架建置 Grafana 外掛頁面,涵蓋場景、面板、變數與下鑽導覽。

Software Development279今天更新

Check Npm

grafana

對 JS/TS 儲存庫的 npm、yarn 或 pnpm 設定進行唯讀供應鏈強化稽核。

Security279今天更新

Mimir

grafana

指導架設與維運 Grafana Mimir,用於可擴充、多租戶、長期的 Prometheus 與 OTLP 指標儲存。

DevOps & Cloud279今天更新

K6 Trend Analysis

grafana

Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for performance regression", "should I tighten thresholds", "are my tests degrading", "show me trends for test X", "analyze my k6 test runs", or "is my test getting slower". Also trigger when a user asks to check all tests in a project -- run this skill once per test and synthesize.

待分類279今天更新