Observability Guidelines

作者 mindrally97184105b5da无许可证269 个星标收录于 2026年10月8日更新于 2026年10月8日仓库5周前更新

Observability guidelines for distributed systems using OpenTelemetry, tracing, metrics, and structured logging

仅含说明DevOps & Cloud
AI 生成的概览

为分布式系统提供 OpenTelemetry 追踪、指标、结构化日志与告警的观测性指南。

功能
该技能为分布式系统和微服务提供一套观测性指南,涵盖 OpenTelemetry 集成、分布式追踪、指标采集、结构化日志、关联与上下文传播、仪表盘与告警、埋点实践,以及生产环境的采样与保留策略。它还列出整洁架构和领域驱动设计等架构模式。它只是纯说明性参考,不生成文件或脚本。
适用场景
在为分布式系统或微服务添加或审查观测能力时使用。适合团队统一追踪、指标、日志和告警规范。也适用于规划生产服务的采样、仪表盘和运行手册。
运行要求
不需要任何工具、软件包或凭据;它仅为说明性内容,不附带脚本。若要落地这些指南,系统需能够使用 OpenTelemetry 以及 OpenTelemetry Collector、Jaeger 或 Prometheus 等相关后端。

Observability Guidelines

Apply these observability principles to ensure comprehensive visibility into distributed systems and microservices.

Core Observability Principles

  • Guide the development of idiomatic, maintainable, and high-performance code with built-in observability
  • Enforce modular design and separation of concerns through Clean Architecture
  • Promote test-driven development and robust observability from the start

OpenTelemetry Integration

  • Use OpenTelemetry for distributed tracing, metrics, and structured logging
  • Start and propagate tracing spans across all service boundaries
  • Use otel.Tracer for creating spans and otel.Meter for collecting metrics
  • Export data to OpenTelemetry Collector, Jaeger, or Prometheus
  • Configure appropriate sampling rates for production environments

Distributed Tracing

  • Trace all incoming requests and propagate context through internal calls
  • Use middleware to instrument HTTP and gRPC endpoints automatically
  • Include trace context in all downstream service calls
  • Create child spans for significant operations within a service
  • Add relevant attributes to spans for debugging and analysis

Metrics Collection

Monitor these key metrics across all services:

  • Request latency: Track p50, p90, p95, and p99 percentiles
  • Throughput: Measure requests per second by endpoint
  • Error rate: Track 4xx and 5xx responses separately
  • Resource usage: Monitor CPU, memory, disk, and network utilization
  • Custom business metrics: Track domain-specific KPIs

Structured Logging

  • Include unique request IDs and trace context in all logs for correlation
  • Use structured logging formats (JSON) for machine parseability
  • Include relevant context: timestamp, service name, trace ID, span ID
  • Log at appropriate levels: DEBUG, INFO, WARN, ERROR
  • Avoid logging sensitive information (PII, credentials)

Architecture Patterns

  • Apply Clean Architecture with handlers, services, repositories, and domain models
  • Use domain-driven design principles for clear boundaries
  • Prioritize interface-driven development with explicit dependency injection
  • Prefer composition over inheritance; favor small, purpose-specific interfaces

Correlation and Context

  • Propagate context through the entire request lifecycle
  • Use correlation IDs for request tracking across services
  • Include service version and deployment information in telemetry
  • Tag traces with relevant business context for filtering
  • Enable trace-to-log and log-to-trace correlation

Alerting and Dashboards

  • Create dashboards for service health and business metrics
  • Set up alerts based on SLOs and error budgets
  • Use anomaly detection for proactive issue identification
  • Document runbooks for common alert scenarios
  • Review and tune alerts regularly to reduce noise

Instrumentation Best Practices

  • Instrument at service boundaries (entry/exit points)
  • Add custom spans for database operations and external calls
  • Include relevant attributes (user ID, request type, etc.)
  • Avoid over-instrumentation that creates noise
  • Use semantic conventions for consistent attribute naming

Production Considerations

  • Configure appropriate sampling rates to balance visibility and cost
  • Use head-based sampling for consistent trace capture
  • Implement tail-based sampling for capturing errors
  • Set retention policies based on debugging needs
  • Monitor observability infrastructure health

来源与署名

来源:mindrally/skills位于observability-guidelines提交9718410

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架