Promql

作者 grafana1ccacf29049fApache-2.0279 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. Covers `rate` vs `irate` vs `increase`, label matchers and regex, `sum / avg / topk / by / without` aggregation, classic + native `histogram_quantile`, ratios with divide-by-zero guards, `absent` / `changes` for staleness, time offsets and `predict_linear`, recording-rule naming, SLO + burn-rate math, and a cardinality-hunting playbook. Use when writing a metric query, fixing wrong p95s, building an error-budget alert, debugging "query is slow", finding the noisy label that blew up cardinality, or migrating a dashboard query to a recording rule — even when the user says "calculate the error rate", "p99 latency", "sum by service", "why is this query slow", or "what's filling Mimir" without naming PromQL.

AI 生成的概览

编写、验证并优化用于 Prometheus、Grafana Mimir 和 Grafana Cloud Metrics 的 PromQL 查询。

功能
提供 PromQL 模式与工作流程,用于编写和验证指标查询,涵盖 rate、irate 与 increase 的区别、标签匹配器、聚合、histogram_quantile、除零保护、陈旧数据处理、时间偏移和 SLO 燃烧率计算。还涉及将缓慢的仪表盘表达式转换为记录规则,以及排查高基数标签。产出查询表达式、记录规则 YAML 以及针对兼容 Prometheus 的查询 API 的验证步骤。
适用场景
适用于编写或修复指标查询、纠正错误的 p95 或 p99 数值、构建错误预算告警、调试缓慢查询、查找导致基数膨胀的标签,或将仪表盘查询迁移为记录规则。
运行要求
需要 Prometheus、Grafana Mimir 或 Grafana Cloud Metrics 查询端点(Grafana Cloud 需基本认证凭据)以及随附的参考模式库。仅为说明文档,不附带脚本。

PromQL Query Patterns

Docs: https://prometheus.io/docs/prometheus/latest/querying/basics/

PromQL returns either an instant vector, a range vector, or a scalar.

Golden rule: rate() / increase() require a range vector ≥ 4× the scrape interval. 60s scrape → use [5m] minimum.

Prerequisites

  • A Prometheus / Mimir / Grafana Cloud endpoint to query (/api/v1/query or via Grafana Explore)
  • The PromQL pattern library in references/patterns.md [blocked]

Common Workflows

1. Write + validate a query

bash
# 0. Point at your Prometheus/Mimir. For Grafana Cloud, use the metrics endpoint#    and add basic auth (-u "<metrics_user>:<token>") to each curl below.PROM=http://localhost:9090   # or https://prometheus-prod-XX.grafana.net/api/prom
# 1. Sketch the query — for "5xx error rate per service":EXPR='sum(rate(http_requests_total{status_code=~"5.."}[5m])) by (service)'
# 2. Validate syntax + that the metric/labels existcurl -sG --data-urlencode "query=${EXPR}" \  "$PROM/api/v1/query" | jq '.status, (.data.result|length)'# Expect: "success" and result count > 0. If 0 — check label spelling and scrape activity:curl -sG --data-urlencode "match[]=http_requests_total" "$PROM/api/v1/series" | jq '.data | length'
# 3. Sanity-check the magnitude — open Grafana Explore, paste the expr,#    confirm the values look right against a known ground truth (k6 run, log count, etc.)

2. Common patterns to copy

Per-status request rate (aggregate AFTER rate):

promql
sum(rate(http_requests_total{job="api"}[5m])) by (status_code)

p95 latency (must keep le in the inner aggregation):

promql
histogram_quantile(0.95,  sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))

Error rate with divide-by-zero guard:

promql
sum(rate(http_requests_total{status_code=~"5.."}[5m]))  / (sum(rate(http_requests_total[5m])) > 0)

Full library (recording rules, SLO burn-rate, offsets, cardinality hunt, native histograms): references/patterns.md [blocked].

3. Convert a slow dashboard query into a recording rule

yaml
# 1. Pick the slow expression, give it a recording-rule namegroups:  - name: http_request_rates    interval: 1m    rules:      - record: job:http_request_duration_p95:rate5m        expr: |          histogram_quantile(0.95,            sum(rate(http_request_duration_seconds_bucket[5m])) by (le, job))
bash
# 2. After rules load, verify the new metric existscurl -sG --data-urlencode "query=job:http_request_duration_p95:rate5m" \  "$PROM/api/v1/query" | jq '.data.result | length'   # → > 0
# 3. Verify it matches the original expression for at least one sample window# (Both queries should produce the same value at the same timestamp.)
# 4. Replace the dashboard panel expression with the recording-rule metric.

Common bugs

  • histogram_quantile returns NaN → forgot by (le) in the inner aggregation
  • "No data" → check the metric exists (/api/v1/series) and the window ≥ 4× scrape interval
  • Wrong rate magnitude → counter was aggregated before rate() (always rate() first)
  • Query timeout → series count too high; use topk(...) + a recording rule + drop high-cardinality labels (see references/patterns.md [blocked])

Resources

来源与署名

来源:grafana/skills位于skills/grafana-core/promql提交1ccacf2

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 grafana/skills 的技能

React 19 Plugin Migration

grafana

指导将 Grafana 插件迁移至 React 19 兼容,按顺序完成构建、依赖与源码修改步骤。

Software Development279今天更新

Plugin Bundle Size

grafana

指导使用 React.lazy、Suspense 和 webpack 代码分割来优化 Grafana 应用插件包体积。

Software Development279今天更新

Grafana Scenes

grafana

使用 @grafana/scenes 框架构建 Grafana 插件页面,涵盖场景、面板、变量与下钻导航。

Software Development279今天更新

Check Npm

grafana

对 JS/TS 仓库的 npm、yarn 或 pnpm 配置进行只读供应链加固审计。

Security279今天更新

Mimir

grafana

指导搭建和运维 Grafana Mimir,用于可扩展、多租户、长期的 Prometheus 与 OTLP 指标存储。

DevOps & Cloud279今天更新

K6 Trend Analysis

grafana

Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for performance regression", "should I tighten thresholds", "are my tests degrading", "show me trends for test X", "analyze my k6 test runs", or "is my test getting slower". Also trigger when a user asks to check all tests in a project -- run this skill once per test and synthesize.

待分类279今天更新