Dd Monitors

作者 datadog-labs5b40c73824ec无许可证177 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Monitor management - list, search, file-based create, and alerting best practices.

仅含说明DevOps & Cloud
AI 生成的概览

使用 pup 管理 Datadog 监控器:列出、搜索、从文件创建、通过停机静默,并遵循告警最佳实践。

功能
该技能提供使用 pup 命令行工具管理 Datadog 监控器的说明。内容涵盖列出和获取监控器、从 JSON 文件创建监控器,以及通过停机负载静默告警。它还记录了告警最佳实践,例如稳定的评估窗口、合理的范围限定、恢复阈值和运行手册链接,并提供将监控器标记为待删除而非直接删除的安全流程。
适用场景
当你需要创建、检查或维护 Datadog 监控器和告警规则时使用。它也适用于审计缺少负责人或告警过多的监控器,以及在计划窗口内静默通知。
运行要求
需要 PATH 中可用的 pup CLI,并通过 pup auth login 完成 Datadog 身份验证。需要访问 Datadog 的网络连接。不附带脚本,仅为说明文档。

Datadog Monitors

Create, manage, and maintain monitors for alerting.

Prerequisites

This requires pup in your path. See Setup Pup.

Command Execution Order (Token-Efficient)

For scoped commands, use this order:

  1. Check context first (prior outputs, conversation, saved values).
  2. If a required value is missing, run a discovery command first.
  3. If still ambiguous, ask the user to confirm.
  4. Then run the target command.
  5. Avoid speculative commands likely to fail.

Quick Start

bash
pup auth login

Common Operations

List Monitors

bash
pup monitors listpup monitors list --tags "team:platform"

Get Monitor

bash
pup monitors get <id>

Create Monitor

bash
pup monitors create --file monitor.json

Silence Alerts (Downtime)

bash
# No pup monitors mute/unmute commands.# Use downtime payloads to silence monitor notifications.pup downtime create --file downtime.jsonpup downtime cancel <downtime_id>

Monitor Creation Best Practices

1. Avoid Alert Fatigue

RuleWhy
No flapping alertsUse last_Xm not last_1m
Meaningful thresholdsBased on SLOs, not guesses
Actionable alertsIf no action needed, don't alert
Include runbook@runbook-url in message
python
# WRONG - will flap constantlyquery = "avg(last_1m):avg:system.cpu.user{*} > 50"  # ❌ Too sensitive
# CORRECT - stable alertingquery = "avg(last_5m):avg:system.cpu.user{env:prod} by {host} > 80"  # ✅ Reasonable window

2. Use Proper Scoping

python
# WRONG - alerts on everythingquery = "avg(last_5m):avg:system.cpu.user{*} > 80"  # ❌ No scope
# CORRECT - scoped to what mattersquery = "avg(last_5m):avg:system.cpu.user{env:prod,service:api} by {host} > 80"  # ✅

3. Set Recovery Thresholds

python
monitor = {    "query": "avg(last_5m):avg:system.cpu.user{env:prod} > 80",    "options": {        "thresholds": {            "critical": 80,            "critical_recovery": 70,  # ✅ Prevents flapping            "warning": 60,            "warning_recovery": 50        }    }}

4. Include Context in Messages

python
message = """## High CPU Alert
Host: {{host.name}}Current Value: {{value}}Threshold: {{threshold}}
### Runbook1. Check top processes: `ssh {{host.name}} 'top -bn1 | head -20'`2. Check recent deploys3. Scale if needed
@slack-ops @pagerduty-oncall"""

NEVER Delete Monitors Directly

Use safe deletion workflow (same as dashboards):

python
def safe_mark_monitor_for_deletion(monitor_id: str, client) -> bool:    """Mark monitor instead of deleting."""    monitor = client.get_monitor(monitor_id)    name = monitor.get("name", "")        if "[MARKED FOR DELETION]" in name:        print(f"Already marked: {name}")        return False        new_name = f"[MARKED FOR DELETION] {name}"    client.update_monitor(monitor_id, {"name": new_name})    print(f"✓ Marked: {new_name}")    return True

Monitor Types

TypeUse Case
metric alertCPU, memory, custom metrics
query alertComplex metric queries
service checkAgent check status
event alertEvent stream patterns
log alertLog pattern matching
compositeCombine multiple monitors
apmAPM metrics

Audit Monitors

bash
# Find monitors without ownerspup monitors list | jq '.[] | select(.tags | contains(["team:"]) | not) | {id, name}'
# Find noisy monitors (high alert count)pup monitors list | jq 'sort_by(.overall_state_modified) | .[:10] | .[] | {id, name, status: .overall_state}'

Downtime vs Muting

UseWhen
DowntimeAny planned silence window
Monitor editQuery/threshold behavior changes
bash
# Downtime (preferred)pup downtime create --file downtime.json

Failure Handling

ProblemFix
Alert not firingCheck query returns data, thresholds
Too many alertsIncrease window, add recovery threshold
No data alertsCheck agent connectivity, metric exists
Auth errorpup auth refresh

References

来源与署名

来源:datadog-labs/agent-skills位于dd-monitors提交5b40c73

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架