dt-alerting
Configure and understand the full alerting lifecycle in Dynatrace — from anomaly detector setup through Grail event storage, problem grouping, and workflow notification delivery.
The Alerting Lifecycle
When to Use This Skill
- Detector setup — "How do I create an anomaly detector?", "What kind of detector should I use?", "What is the difference between adaptive and seasonal?"
- Alert event history — "Query all alert events for this service", "Show me which metrics triggered alerts last week"
- Problem denoising — "Why did these two alerts merge into one problem?", "How does Davis group alerts?"
- Notification setup — "How do I send a Slack message when a problem opens?", "Set up a ServiceNow ticket on critical problems"
- Best practices — "How do I avoid alert storms?", "Which sensitivity setting should I use?"
- Over-alerting analysis — "Why am I getting too many alerts?", "How do I reduce alert fatigue?", "Which detector is firing the most?", "How do I tune sensitivity or thresholds to avoid noise?"
- Notification routing — "How do I route alerts to the right team?", "Set up scalable problem filters in workflows", "Send Slack notifications only to the team responsible for the affected service"
Agent Instructions
First step for any alerting setup request — Before recommending a specific
detector or model, load references/anomaly-detectors.md and use its category
and model decision guide to identify which detector category (DQL-based, Edge,
Pipeline, Synthetic, External) and which model (Static, Adaptive, Seasonal)
best fits the user's use-case. Only proceed with configuration guidance once
the right detector type has been established.
Consolidate, don't multiply — When a user asks to alert on multiple
entities of the same kind (e.g. "alert on services A, B, and C"), always
recommend a single combined detector rather than one detector per entity.
Use by: { <dimension> } in the DQL timeseries call to split results per
entity, and a single filter: clause to scope to the relevant entities.
Pair the combined detector with a single dt.alert_group tag shared
across all alert conditions and the corresponding workflow notification filter.
This keeps the number of detector configs small, ensures consistent routing,
and makes the workflow notification channel reusable for future entities added
to the same group.
Example for three services — one detector, one workflow:
Set dt.alert_group: "checkout-team" in the detector's event properties, then
filter the notification workflow on matchesPhrase(dt.alert_group, "checkout-team").
If a new service must be covered, add it to the single filter: list — no new
detector or workflow rule needed.
Intent Mapping
Analyzing existing problems — If the user wants to query or investigate active/closed problems (root cause, impact, trending), load
dt-obs-problemsinstead. This skill covers configuration and flow, not problem query analytics.
Detector health monitoring — If the user asks whether detectors are running or failing, load
dt-platform(ANALYZER_EXECUTION_EVENT, ANOMALY_DETECTOR_STATUS_EVENT). This skill covers setup, not operational health.
Prerequisites
- Access to a Dynatrace environment with Settings v2 write permissions for detector configuration
- For querying alert history: DQL permissions on
dt.davis.events - Load
dt-dql-essentialsbefore writing DQL queries
Knowledge Base Structure
Key Concepts
Alert Source Categories
Five fundamental categories of anomaly detectors, distinguished by where detection runs and how the alert event reaches Dynatrace:
See references/anomaly-detectors.md for the full breakdown of each category,
including trade-offs and configuration entry points.
Detector Models at a Glance
Davis Events vs. Problems
A single problem typically contains multiple events. Querying problems gives the operational view; querying events gives the raw alert history.
Problem Denoising
For questions about why alerts merged into a problem or how Davis groups
events, load dt-obs-problems — the merge logic and rules are documented in
dt-obs-problems/references/problem-merging.md. This skill covers alert
configuration and flow only.
Quick Start
Check What Alerts Fired in the Last 24 Hours
Check Alert Volume by Category
See All Active Problems (→ load dt-obs-problems for full query patterns)
Best Practices
- Match the model to the metric's behavior — Use static for hard SLO boundaries, adaptive for metrics without a natural fixed limit, seasonal for anything that follows business hours or weekly patterns.
- Scope detectors narrowly — An entity selector that covers only relevant entities reduces noise and makes problems more actionable.
- Tune sensitivity before going to production — Start with LOW sensitivity and move to MEDIUM or HIGH only after observing false-positive rates.
- Let Davis denoise before notifying — Trigger workflow notifications on problems, not individual alert events. A problem groups correlated alerts so you notify once per incident, not once per metric.
- Filter notifications by severity level — Route
event.severity <= 2problems to on-call channels immediately; routeevent.severity >= 3problems to lower- urgency channels. Either set severity in the detector config or assign in a pipeline rule or workflow. - Use
dt.alert_groupevent property for routing — Assigndt.alert_groupto route alerts to the right team. Either set a static value in the detector config, use dynamic assignment through DQL query result mapping or assign in a pipeline rule. - Combine same-condition alerts into one detector and one workflow — When
alerting on multiple entities with the same metric and threshold, merge them
into a single DQL-based detector using
by: { <dimension> }and a combinedfilter:clause. Assign the samedt.alert_groupvalue to every condition in that detector and point the workflow notification channel at that single group. One detector + one workflow per logical alert group scales better than N detectors + N notification rules, and adding a new entity is a one-line filter change rather than a full detector/workflow addition.
Related Skills
- dt-obs-problems — Querying, analyzing, and trending detected problems
- dt-obs-predictive-analytics — Ad-hoc anomaly and novelty detection using MCP analyzer tools (not persistent alert configs)
- dt-platform — Operational health of anomaly detectors (execution events, failure rates)
- dt-platform-costs — Query costs generated by anomaly detector DQL
- dt-sdlc-quality-gates — Site Reliability Guardian for deployment gate alerting
- dt-dql-essentials — DQL syntax for writing detector queries and alert history queries


