Dd Monitors

作者 datadog-labs5b40c73824ec無授權條款177 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Monitor management - list, search, file-based create, and alerting best practices.

僅含說明DevOps & Cloud
AI 產生的概覽

使用 pup 管理 Datadog 監控器:列出、搜尋、從檔案建立、透過停機靜默,並遵循告警最佳實務。

功能
此技能提供使用 pup 命令列工具管理 Datadog 監控器的說明。內容涵蓋列出與取得監控器、從 JSON 檔案建立監控器,以及透過停機承載內容靜默告警。它也記錄了告警最佳實務,例如穩定的評估視窗、適當的範圍限定、恢復閾值和執行手冊連結,並提供將監控器標記為待刪除而非直接刪除的安全流程。
適用情境
當你需要建立、檢查或維護 Datadog 監控器與告警規則時使用。它也適用於稽核缺少負責人或告警過多的監控器,以及在計畫視窗內靜默通知。
執行需求
需要 PATH 中可用的 pup CLI,並透過 pup auth login 完成 Datadog 身分驗證。需要存取 Datadog 的網路連線。不附帶指令碼,僅為說明文件。

Datadog Monitors

Create, manage, and maintain monitors for alerting.

Prerequisites

This requires pup in your path. See Setup Pup.

Command Execution Order (Token-Efficient)

For scoped commands, use this order:

  1. Check context first (prior outputs, conversation, saved values).
  2. If a required value is missing, run a discovery command first.
  3. If still ambiguous, ask the user to confirm.
  4. Then run the target command.
  5. Avoid speculative commands likely to fail.

Quick Start

bash
pup auth login

Common Operations

List Monitors

bash
pup monitors listpup monitors list --tags "team:platform"

Get Monitor

bash
pup monitors get <id>

Create Monitor

bash
pup monitors create --file monitor.json

Silence Alerts (Downtime)

bash
# No pup monitors mute/unmute commands.# Use downtime payloads to silence monitor notifications.pup downtime create --file downtime.jsonpup downtime cancel <downtime_id>

Monitor Creation Best Practices

1. Avoid Alert Fatigue

RuleWhy
No flapping alertsUse last_Xm not last_1m
Meaningful thresholdsBased on SLOs, not guesses
Actionable alertsIf no action needed, don't alert
Include runbook@runbook-url in message
python
# WRONG - will flap constantlyquery = "avg(last_1m):avg:system.cpu.user{*} > 50"  # ❌ Too sensitive
# CORRECT - stable alertingquery = "avg(last_5m):avg:system.cpu.user{env:prod} by {host} > 80"  # ✅ Reasonable window

2. Use Proper Scoping

python
# WRONG - alerts on everythingquery = "avg(last_5m):avg:system.cpu.user{*} > 80"  # ❌ No scope
# CORRECT - scoped to what mattersquery = "avg(last_5m):avg:system.cpu.user{env:prod,service:api} by {host} > 80"  # ✅

3. Set Recovery Thresholds

python
monitor = {    "query": "avg(last_5m):avg:system.cpu.user{env:prod} > 80",    "options": {        "thresholds": {            "critical": 80,            "critical_recovery": 70,  # ✅ Prevents flapping            "warning": 60,            "warning_recovery": 50        }    }}

4. Include Context in Messages

python
message = """## High CPU Alert
Host: {{host.name}}Current Value: {{value}}Threshold: {{threshold}}
### Runbook1. Check top processes: `ssh {{host.name}} 'top -bn1 | head -20'`2. Check recent deploys3. Scale if needed
@slack-ops @pagerduty-oncall"""

NEVER Delete Monitors Directly

Use safe deletion workflow (same as dashboards):

python
def safe_mark_monitor_for_deletion(monitor_id: str, client) -> bool:    """Mark monitor instead of deleting."""    monitor = client.get_monitor(monitor_id)    name = monitor.get("name", "")        if "[MARKED FOR DELETION]" in name:        print(f"Already marked: {name}")        return False        new_name = f"[MARKED FOR DELETION] {name}"    client.update_monitor(monitor_id, {"name": new_name})    print(f"✓ Marked: {new_name}")    return True

Monitor Types

TypeUse Case
metric alertCPU, memory, custom metrics
query alertComplex metric queries
service checkAgent check status
event alertEvent stream patterns
log alertLog pattern matching
compositeCombine multiple monitors
apmAPM metrics

Audit Monitors

bash
# Find monitors without ownerspup monitors list | jq '.[] | select(.tags | contains(["team:"]) | not) | {id, name}'
# Find noisy monitors (high alert count)pup monitors list | jq 'sort_by(.overall_state_modified) | .[:10] | .[] | {id, name, status: .overall_state}'

Downtime vs Muting

UseWhen
DowntimeAny planned silence window
Monitor editQuery/threshold behavior changes
bash
# Downtime (preferred)pup downtime create --file downtime.json

Failure Handling

ProblemFix
Alert not firingCheck query returns data, thresholds
Too many alertsIncrease window, add recovery threshold
No data alertsCheck agent connectivity, metric exists
Auth errorpup auth refresh

References

來源與署名

來源:datadog-labs/agent-skills位於dd-monitors提交5b40c73

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架