Dd Monitors

by datadog-labs5b40c73824ecNo license177 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Monitor management - list, search, file-based create, and alerting best practices.

Instructions onlyDevOps & Cloud
AI-generated overview

Manage Datadog monitors with pup: list, search, create from files, silence via downtime, and follow alerting best practices.

What it does
This skill provides instructions for managing Datadog monitors using the pup command-line tool. It covers listing and getting monitors, creating monitors from JSON files, and silencing alerts through downtime payloads. It also documents alerting best practices such as stable evaluation windows, proper scoping, recovery thresholds, and runbook links, plus a safe workflow that marks monitors for deletion instead of deleting them.
When to use it
Use it when you need to create, inspect, or maintain Datadog monitors and alerting rules. It is also useful for auditing monitors for missing owners or noisy alerts, and for silencing notifications during planned windows.
Requirements
Requires the pup CLI available in PATH and Datadog authentication via pup auth login. Network access to Datadog is needed. No scripts are shipped; the skill is instructions only.

Datadog Monitors

Create, manage, and maintain monitors for alerting.

Prerequisites

This requires pup in your path. See Setup Pup.

Command Execution Order (Token-Efficient)

For scoped commands, use this order:

  1. Check context first (prior outputs, conversation, saved values).
  2. If a required value is missing, run a discovery command first.
  3. If still ambiguous, ask the user to confirm.
  4. Then run the target command.
  5. Avoid speculative commands likely to fail.

Quick Start

bash
pup auth login

Common Operations

List Monitors

bash
pup monitors listpup monitors list --tags "team:platform"

Get Monitor

bash
pup monitors get <id>

Create Monitor

bash
pup monitors create --file monitor.json

Silence Alerts (Downtime)

bash
# No pup monitors mute/unmute commands.# Use downtime payloads to silence monitor notifications.pup downtime create --file downtime.jsonpup downtime cancel <downtime_id>

Monitor Creation Best Practices

1. Avoid Alert Fatigue

RuleWhy
No flapping alertsUse last_Xm not last_1m
Meaningful thresholdsBased on SLOs, not guesses
Actionable alertsIf no action needed, don't alert
Include runbook@runbook-url in message
python
# WRONG - will flap constantlyquery = "avg(last_1m):avg:system.cpu.user{*} > 50"  # ❌ Too sensitive
# CORRECT - stable alertingquery = "avg(last_5m):avg:system.cpu.user{env:prod} by {host} > 80"  # ✅ Reasonable window

2. Use Proper Scoping

python
# WRONG - alerts on everythingquery = "avg(last_5m):avg:system.cpu.user{*} > 80"  # ❌ No scope
# CORRECT - scoped to what mattersquery = "avg(last_5m):avg:system.cpu.user{env:prod,service:api} by {host} > 80"  # ✅

3. Set Recovery Thresholds

python
monitor = {    "query": "avg(last_5m):avg:system.cpu.user{env:prod} > 80",    "options": {        "thresholds": {            "critical": 80,            "critical_recovery": 70,  # ✅ Prevents flapping            "warning": 60,            "warning_recovery": 50        }    }}

4. Include Context in Messages

python
message = """## High CPU Alert
Host: {{host.name}}Current Value: {{value}}Threshold: {{threshold}}
### Runbook1. Check top processes: `ssh {{host.name}} 'top -bn1 | head -20'`2. Check recent deploys3. Scale if needed
@slack-ops @pagerduty-oncall"""

NEVER Delete Monitors Directly

Use safe deletion workflow (same as dashboards):

python
def safe_mark_monitor_for_deletion(monitor_id: str, client) -> bool:    """Mark monitor instead of deleting."""    monitor = client.get_monitor(monitor_id)    name = monitor.get("name", "")        if "[MARKED FOR DELETION]" in name:        print(f"Already marked: {name}")        return False        new_name = f"[MARKED FOR DELETION] {name}"    client.update_monitor(monitor_id, {"name": new_name})    print(f"✓ Marked: {new_name}")    return True

Monitor Types

TypeUse Case
metric alertCPU, memory, custom metrics
query alertComplex metric queries
service checkAgent check status
event alertEvent stream patterns
log alertLog pattern matching
compositeCombine multiple monitors
apmAPM metrics

Audit Monitors

bash
# Find monitors without ownerspup monitors list | jq '.[] | select(.tags | contains(["team:"]) | not) | {id, name}'
# Find noisy monitors (high alert count)pup monitors list | jq 'sort_by(.overall_state_modified) | .[:10] | .[] | {id, name, status: .overall_state}'

Downtime vs Muting

UseWhen
DowntimeAny planned silence window
Monitor editQuery/threshold behavior changes
bash
# Downtime (preferred)pup downtime create --file downtime.json

Failure Handling

ProblemFix
Alert not firingCheck query returns data, thresholds
Too many alertsIncrease window, add recovery threshold
No data alertsCheck agent connectivity, metric exists
Auth errorpup auth refresh

References

Source and attribution

Source:datadog-labs/agent-skillsindd-monitorsat commit5b40c73

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal