Distributed Tracing

作者 wshobson46891e7e60da無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks. Use when debugging microservices, analyzing request flows, or implementing observability for distributed systems.

僅含說明DevOps & Cloud
AI 產生的概覽

指導使用 Jaeger 和 Tempo 實作分散式追蹤,以追蹤微服務之間的請求。

功能
此技能提供使用 Jaeger 和 Tempo 為微服務實作分散式追蹤的指引。內容涵蓋取樣、span 標籤、情境傳遞、關聯日誌,以及排解追蹤缺失或延遲開銷問題。它也指向一份包含詳細模式與範例的參考檔案。
適用情境
適用於偵錯微服務中的延遲或錯誤、梳理服務相依關係,或為分散式系統增加可觀測性。也適合分析請求路徑並找出效能瓶頸。
執行需求
不包含指令碼,僅為說明性內容。實作時需要可用的追蹤工具(如 Jaeger 或 Tempo);日誌範例還涉及 OpenTelemetry 與 Python。

Distributed Tracing

Implement distributed tracing with Jaeger and Tempo for request flow visibility across microservices.

Purpose

Track requests across distributed systems to understand latency, dependencies, and failure points.

When to Use

  • Debug latency issues
  • Understand service dependencies
  • Identify bottlenecks
  • Trace error propagation
  • Analyze request paths

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Best Practices

  1. Sample appropriately (1-10% in production)
  2. Add meaningful tags (user_id, request_id)
  3. Propagate context across all service boundaries
  4. Log exceptions in spans
  5. Use consistent naming for operations
  6. Monitor tracing overhead (<1% CPU impact)
  7. Set up alerts for trace errors
  8. Implement distributed context (baggage)
  9. Use span events for important milestones
  10. Document instrumentation standards

Integration with Logging

Correlated Logs

python
import loggingfrom opentelemetry import trace
logger = logging.getLogger(__name__)
def process_request():    span = trace.get_current_span()    trace_id = span.get_span_context().trace_id
    logger.info(        "Processing request",        extra={"trace_id": format(trace_id, '032x')}    )

Troubleshooting

No traces appearing:

  • Check collector endpoint
  • Verify network connectivity
  • Check sampling configuration
  • Review application logs

High latency overhead:

  • Reduce sampling rate
  • Use batch span processor
  • Check exporter configuration

Related Skills

  • prometheus-configuration - For metrics
  • grafana-dashboards - For visualization
  • slo-implementation - For latency SLOs

來源與署名

來源:wshobson/agents位於plugins/observability-monitoring/skills/distributed-tracing提交46891e7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架