Slos And Triggers

作者 honeycombiob169d7d1c76a無授權條款27 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫2 週前更新

Decision heuristics for interpreting Honeycomb SLO compliance, budget burn rates, and trigger status — what the numbers mean and what action to take, including detecting misconfigured SLIs, deciding when to freeze deploys vs page on-call, and designing burn alert thresholds. Load this skill before calling get_slos or get_triggers. Trigger phrases: "check our SLOs", "are we meeting our SLOs", "which SLOs are healthy", "is the error budget OK", "are any alerts firing", "what's the burn rate", "set up an SLO", "create a trigger", "configure alerts", "set up burn alerts", "check trigger status", "starting on-call", "reliability picture", "should we freeze deploys", "is this SLO misconfigured", "are we within budget", "SLO is broken", "budget is negative", or any request about service level objectives, error budgets, burn rates, or alerting in Honeycomb.

僅含說明DevOps & Cloud
AI 產生的概覽

指導解讀 Honeycomb SLO 達標情形、錯誤預算消耗速率與觸發器狀態,並設計 SLO 與告警。

功能
此技能提供解讀 Honeycomb 可靠性資料的決策準則:如何理解 SLO 剩餘預算、消耗速率與觸發器狀態,以及每種狀態對應什麼行動。內容涵蓋 SLI 與目標的設計、SLO 與觸發器之間的取捨、耗盡告警與消耗速率告警的設定,以及多服務 SLO。它也會指出可能設定錯誤的 SLI,並建議在建立 SLO 或觸發器前先與使用者確認參數。產出是指引與建議,而非檔案或程式碼。
適用情境
適用於檢查服務是否達成 SLO、查看錯誤預算或消耗速率,或判斷是否有告警觸發時。也用於建立新的 SLO、觸發器或消耗告警,以及決定是否凍結部署或呼叫值班人員。
執行需求
需要 Honeycomb 的 get_slos 與 get_triggers 工具以及 Honeycomb 工作區;SLO 需要 Pro 或 Enterprise 方案。此技能不含指令碼,只有參考文件。

Honeycomb SLOs and Triggers

Guidance for configuring and reasoning about reliability in Honeycomb. The get_slos and get_triggers tools document their own parameters — this skill focuses on designing effective SLOs, choosing between SLOs and triggers, and interpreting what the numbers mean.

Availability: SLOs require Pro or Enterprise plan. Triggers available on all plans.

SLO vs Trigger — When to Use Which

QuestionSLOTrigger
"Are we meeting our reliability commitments?"YesNo
"Is something broken right now?"NoYes
"How fast are we burning our error budget?"Yes (burn alerts)No
"Did error count exceed a threshold?"NoYes
"Should we slow down deploys?"Yes (budget remaining)No

Rule of thumb: SLOs measure reliability against commitments over time. Triggers catch immediate operational issues.

Designing Effective SLOs

Define the SLI

An SLI is a per-event boolean: was this event successful? Implemented as a calculated field returning undefined (not a relevant event), 1 (success), or 0 (failure).

  • Format: IF(<qualifying-condition>, <success-condition>) The qualifying condition filters to relevant events; the success condition defines what counts as success. If the qualifying condition is not met, the formula returns undefined, and the SLI is unpopulated.
  • Specific Qualifying Condition: Choose the relevant subset of events (e.g. AND(EQUALS($http.route, "/checkout"), NOT(EXISTS($trace.parent_id))) for root spans of checkout endpoint)
  • Latency Success Condition: LTE(duration_ms, 500) — requests faster than 500ms
  • Availability Success Condition: LTE(http.status_code, 499) — non-5xx responses
  • Business Logic Success Condition: EQUALS(checkout.status, "completed") — successful checkouts

Set the Target

  • Start conservative (99% before 99.99%)
  • Measure current baseline first with P50/P99 queries
  • Set target slightly above current performance
  • Ask: what reliability do users actually need?

Configure Exhaustion Time Alerts

At minimum, two alerts:

  • Near exhaustion (exhaustion time ~4h): pages on-call via PagerDuty
  • Trending to exhaustion (budget rate over 24h): notifies team via Slack

Configure Burn Rate Alerts

Detect fast burns even if the budget isn't close to exhaustion yet. For example:

  • 1h burn rate > 10x — page on-call

Recommend these alerts to the user after creating the SLO. Agents do not have the ability to set up these alerts or their recipients.

Best Practices

  • Measure close to the user (at the edge, not deep in the stack)
  • Design around user workflows, not team boundaries
  • Favor broad SLOs over many narrow ones
  • Start with one SLO, reduce noise, then expand

Interpreting SLO Status

When reviewing SLOs with get_slos:

  • Budget remaining > 50%: Healthy — room for risk
  • Budget remaining 10-50%: Caution — slow down changes
  • Budget remaining < 10%: At risk — freeze non-critical deploys
  • Budget negative: Breached — investigate immediately with the production-investigation skill
  • Compliance at 0%: Likely misconfigured SLI (wrong column, inverted logic, no matching events) — check the SLI definition

Configuring Triggers

Prefer Count-Based Over Percentile-Based

"50 requests slower than 2s" is more actionable than "P99 is 2100ms." Use COUNT WHERE duration_ms > threshold instead of P99 triggers.

Common Patterns

  • Error spike: COUNT WHERE error = true, threshold > N in 5 min
  • Slow requests: COUNT WHERE duration_ms > 2000, threshold > N in 5 min
  • Traffic drop: COUNT WHERE is_root, threshold < N in 10 min (below normal)

Best Practices

  • Name: What the alert is. Description: What to do (link to runbook).
  • Set duration 5-10 min minimum to avoid flapping
  • Start less sensitive, tighten based on false positive rate

Multi-Service SLOs

Share a single error budget across up to 10 services.

  • SLI must be an environment-level calculated field
  • Events from included services weighted equally
  • Use cases: multiple edge services, monolith-to-microservices migration

Check in with the user

Workspaces in Honeycomb have a limited number of SLOs and triggers. Before executing the create tool, check in with the user. Display all parameters and your reasoning, and ask for confirmation.

Constructing links to SLOs

The tools you have will not let you link directly to the SLO page in Honeycomb. Instead, you can link to the list of SLOs.

/<team_slug>/environments/<environment_slug>/slos

Additional Resources

Reference Files

  • ${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/slo-design-guide.md — Detailed SLO design methodology, multi-service SLOs, error budget math
  • ${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/trigger-examples.md — Complete trigger example library organized by use case
  • ${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/alerting-strategy.md — How to combine SLO burn alerts and triggers into a cohesive alerting strategy

Cross-References

  • For constructing SLI queries and calculated fields, see the query-patterns skill
  • For investigating SLO budget burn, see the production-investigation skill

來源與署名

來源:honeycombio/agent-skill位於honeycomb/skills/slos-and-triggers提交b169d7d

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 honeycombio/agent-skill 的技能

Verify Recent Trace

honeycombio

查詢 Honeycomb,找出近期測試產生的 trace 並回報結果。

DevOps & Cloud272 週前更新

Production Investigation

honeycombio

在 Honeycomb 中調查正式環境問題並找出根因的結構化工作流程。

DevOps & Cloud272 週前更新

Observability Fundamentals

honeycombio

First principles behind observability — wide events, high cardinality, the core analysis loop, events vs metrics vs logs, and how instrumentation connects to debugging outcomes. Grounds recommendations in first principles rather than tool-specific how-to. Trigger phrases: "what is observability", "why observability", "why Honeycomb", "events vs metrics vs logs", "events vs metrics", "events vs logs", "metrics vs logs", "why wide events", "what is high cardinality", "core analysis loop", "observability vs monitoring", "what is dimensionality", "explain observability", or any conceptual question about observability or why Honeycomb's approach differs from traditional monitoring.

待分類272 週前更新

Metrics Queries

honeycombio

說明如何在 Honeycomb 中正確查詢 OpenTelemetry 指標資料集,涵蓋允許的操作、時間聚合與直方圖。

Data & Analytics272 週前更新

Create Honeycomb Board

honeycombio

透過 MCP 工具設計並建立包含查詢、SLO 與文字面板的 Honeycomb 看板(儀表板)。

DevOps & Cloud272 週前更新

Beeline Migration

honeycombio

指導從 Honeycomb Beelines 分兩階段遷移到 OpenTelemetry 埋點。

DevOps & Cloud272 週前更新