Dd Audit Cost Spike Investigation

作者 datadog-labs5b40c73824ec无许可证177 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Investigate a Datadog product usage or cost spike by correlating Usage Metering data (when/what spiked) with Audit Trail config changes (who changed what in the preceding window).

AI 生成的概览

将 Datadog 用量计量数据与 Audit Trail 配置变更关联,以解释成本或用量激增。

功能
指导一项调查:先查询 Datadog 每小时用量计量数据,确定激增发生的时间及所属产品族,再检索此前时间窗口内的 Audit Trail 配置变更。它会将结果收窄到与激增产品相关的审计类别,并输出报告,指出可能的因果变更、置信度以及后续步骤。它识别的是配置变更,而非提交数据的具体用户或进程。
适用场景
当某个 Datadog 产品的用量或成本意外上升,需要找出背后的配置变更时使用。适用于需要将账单数据与变更历史关联起来的 FinOps 或运维调查。
运行要求
需要 Datadog 访问权限:通过 pup CLI 进行 OAuth2 登录以查询审计日志,并提供 DD_API_KEY、DD_APP_KEY 和 DD_SITE 以查询用量计量数据。需要访问 Datadog API 的网络连接,以及 pup CLI 和 jq。该技能不附带脚本,仅为操作说明。

Audit Trail: Cost / Usage Spike Investigation

Identify what caused a Datadog usage spike by correlating billing data with configuration change history.

The causal chain is: someone changed something → that change increased data volume → usage spiked → cost went up. Usage Metering tells you when and what; Audit Trail tells you who made the change.

Prerequisites

bash
pup auth login   # OAuth2 (recommended) — covers audit queries# Usage Metering queries also need DD_API_KEY + DD_APP_KEYexport DD_API_KEY=<your-api-key>export DD_APP_KEY=<your-app-key>export DD_SITE=datadoghq.com

Scope Boundary

This skill identifies configuration changes that may have caused a spike. It does not identify which specific user or process submitted the data (e.g., which service sent the LLM spans). For per-submission attribution, use LLM Observability traces or APM instrumentation.

Investigation Workflow

Step 1 — Identify the spike window and product family

bash
START=$(date -u -v-7d +"%Y-%m-%dT%H:%M:%SZ" 2>/dev/null || date -u -d "7 days ago" +"%Y-%m-%dT%H:%M:%SZ")END=$(date -u +"%Y-%m-%dT%H:%M:%SZ")
curl -s -G "https://api.${DD_SITE}/api/v2/usage/hourly_usage" \  -H "DD-API-KEY: ${DD_API_KEY}" \  -H "DD-APPLICATION-KEY: ${DD_APP_KEY}" \  --data-urlencode "filter[timestamp][start]=${START}" \  --data-urlencode "filter[timestamp][end]=${END}" \  --data-urlencode "filter[product_families]=all" \  | jq '[.data[] | {      timestamp: .attributes.timestamp,      product: .attributes.product_family,      measurements: [.attributes.measurements[] | {type: .usage_type, value: .value}]    }]'

Product families with LLM/AI coverage: llm_observability, bits_ai, logs, apm

Step 2 — Pinpoint the spike

From Step 1, identify the hour/day where volume jumped. Note the timestamp as SPIKE_TIME.

Step 3 — Search Audit Trail for config changes in the 24h preceding the spike

bash
pup audit-logs search \  --query "@action:(created OR modified OR deleted)" \  --from "SPIKE_TIME_MINUS_24H" \  --to "SPIKE_TIME" \  --limit 200 \  -o json \  | jq '[.data[] | {      timestamp: .attributes.timestamp,      user: .attributes.attributes.usr.email,      actor_type: .attributes.attributes.evt.actor.type,      action: .attributes.attributes.action,      event_category: .attributes.attributes.evt.name,      resource_type: .attributes.attributes.asset.type,      resource_id: .attributes.attributes.asset.id    }]'

Note: --from and --to accept ISO timestamps (e.g., 2026-05-01T14:00:00Z) or relative values (1h, 24h, 7d).

Step 4 — Narrow to product-relevant config changes

Filter to the audit categories most likely to affect the spiking product:

If this product spikedAdd to query
llm_observability@evt.name:(Integration OR APM OR "Log Management")
logs / indexed_logs@evt.name:"Log Management" @asset.type:(pipeline OR index OR exclusion_filter)
apm / indexed_spans@evt.name:APM @asset.type:(retention_filter OR sampling_rate)
rum@evt.name:RUM
metrics@evt.name:Metrics

Example for LLM Observability spike:

bash
pup audit-logs search \  --query "@evt.name:(Integration OR APM OR \"Log Management\") @action:(created OR modified)" \  --from "SPIKE_TIME_MINUS_24H" \  --to "SPIKE_TIME" \  --limit 100 \  -o json \  | jq '[.data[] | {      timestamp: .attributes.timestamp,      user: .attributes.attributes.usr.email,      action: .attributes.attributes.action,      category: .attributes.attributes.evt.name,      resource_type: .attributes.attributes.asset.type,      resource_id: .attributes.attributes.asset.id    }]'

Output Format

Usage spike detected:  Product: <product_family>  Spike time: <SPIKE_TIME>  Volume: <baseline> → <spike_value> (<magnitude>×)
Configuration changes in 24h preceding spike:  <timestamp> | <user_email> | <action> <resource_type> <resource_id> | <category>
Likely causal change: <most-proximate change matching the product family>
Confidence: HIGH (single clear change) / MEDIUM (multiple candidates) / LOW (no matching changes)
Next steps:  - Confirm with <user_email> whether the change was intentional  - If unintentional: revert <resource_id> and monitor volume  - If intentional: update cost forecasts and alert thresholds

When No Causal Change Is Found

  1. The change may predate the 24h window — expand to 72h
  2. The increase may be from application-side instrumentation changes — check deploys
  3. The increase may be organic traffic growth — correlate with product launch or traffic event

References

来源与署名

来源:datadog-labs/agent-skills位于dd-audit/cost-spike-investigation提交5b40c73

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架