Monitoring Guidelines

作者 mindrally97184105b5da無授權條款269 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫5 週前更新

Monitoring guidelines for applications and infrastructure including metrics collection, alerting strategies, and SLO-based monitoring

僅含說明DevOps & Cloud
AI 產生的概覽

適用於應用程式與基礎架構的監控指引:指標、告警、SLO、儀表板、健康檢查與容量規劃。

功能
提供一套針對應用程式與基礎架構的監控原則參考。內容涵蓋應用程式、基礎架構與業務各層應留意的指標,告警設計原則與以 SLO 為基礎的告警(含錯誤預算與燃燒率),以及儀表板設計、資料收集與保留、健康檢查與探針、事件應變和容量規劃。產出為指引文字,而非產生的檔案或程式碼。
適用情境
適用於建置或檢視服務與基礎架構的監控和告警、定義 SLO 與告警門檻、設計儀表板,或規劃健康檢查、事件應變與容量時使用。適合需要一份可套用的監控實務清單的團隊。
執行需求
不需要任何工具、套件、執行環境、憑證或網路存取;僅為說明性指引,不附帶指令碼。

Monitoring Guidelines

Apply these monitoring principles to ensure system reliability, performance visibility, and proactive issue detection.

Core Monitoring Principles

  • Monitor the four golden signals: latency, traffic, errors, and saturation
  • Implement monitoring as code for reproducibility
  • Design monitoring around user experience and business impact
  • Use SLOs (Service Level Objectives) to guide alerting decisions
  • Balance comprehensive coverage with actionable insights

Key Metrics to Monitor

Application Metrics

  • Request rate (requests per second)
  • Error rate (percentage of failed requests)
  • Response time (p50, p90, p95, p99 latencies)
  • Active connections and concurrent users
  • Queue depths and processing times

Infrastructure Metrics

  • CPU utilization and load average
  • Memory usage and available memory
  • Disk I/O and available storage
  • Network throughput and error rates
  • Container and pod health (for Kubernetes)

Business Metrics

  • Transaction volumes and values
  • User signups and conversions
  • Feature usage and adoption rates
  • Revenue-impacting events
  • Customer satisfaction indicators

Alerting Strategy

Alert Design Principles

  • Alert on symptoms, not causes
  • Make alerts actionable with clear remediation steps
  • Set appropriate severity levels (critical, warning, info)
  • Avoid alert fatigue through proper threshold tuning
  • Include runbook links in alert notifications

SLO-Based Alerting

  • Define SLOs for critical user journeys
  • Calculate error budgets and burn rates
  • Alert when error budget consumption is high
  • Use multi-window, multi-burn-rate alerts
  • Review and adjust SLOs quarterly

Alert Configuration

  • Set meaningful thresholds based on baseline data
  • Use hysteresis to prevent flapping alerts
  • Implement alert dependencies to reduce noise
  • Route alerts to appropriate teams
  • Configure escalation policies

Dashboard Design

Effective Dashboards

  • Create overview dashboards for service health
  • Build detailed dashboards for debugging
  • Use consistent layouts and naming conventions
  • Include time range selectors and drill-down capabilities
  • Display SLO status prominently

Dashboard Content

  • Show current state and recent trends
  • Include comparison to baseline or previous periods
  • Display deployment markers for correlation
  • Add annotations for significant events
  • Include links to related dashboards and logs

Monitoring Tools Integration

Data Collection

  • Use agents or sidecars for metric collection
  • Implement service discovery for dynamic environments
  • Configure appropriate scrape intervals
  • Use push vs pull based on use case
  • Ensure metric cardinality is manageable

Data Storage and Retention

  • Set retention periods based on use case
  • Implement downsampling for long-term storage
  • Use appropriate storage backends for scale
  • Plan for disaster recovery of monitoring data
  • Monitor your monitoring infrastructure

Health Checks and Probes

  • Implement liveness probes for crash detection
  • Use readiness probes for traffic management
  • Create deep health checks that verify dependencies
  • Expose health endpoints in a standard format
  • Monitor health check latency as a metric

Incident Response

  • Use monitoring data to detect incidents early
  • Correlate metrics, logs, and traces during investigation
  • Document findings and update monitoring post-incident
  • Track MTTR (Mean Time to Recovery) metrics
  • Conduct regular monitoring reviews and improvements

Capacity Planning

  • Track resource utilization trends
  • Set alerts for approaching capacity limits
  • Use forecasting for proactive scaling
  • Document capacity requirements and headroom
  • Review capacity quarterly

來源與署名

來源:mindrally/skills位於monitoring-guidelines提交9718410

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架