Distributed Tracing

作者 wshobson46891e7e60da无许可证收录于 2026年10月8日更新于 2026年10月8日

Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks. Use when debugging microservices, analyzing request flows, or implementing observability for distributed systems.

仅含说明DevOps & Cloud
AI 生成的概览

指导使用 Jaeger 和 Tempo 实现分布式追踪,以跟踪微服务间的请求。

功能
该技能提供使用 Jaeger 和 Tempo 为微服务实现分布式追踪的指导。内容涵盖采样、span 标签、上下文传播、关联日志,以及排查追踪缺失或延迟开销问题。它还指向一个包含详细模式与示例的参考文件。
适用场景
适用于调试微服务中的延迟或错误、梳理服务依赖关系,或为分布式系统增加可观测性。也适合分析请求路径并定位性能瓶颈。
运行要求
不包含脚本,仅为说明性内容。实践时需要可用的追踪工具(如 Jaeger 或 Tempo);日志示例还涉及 OpenTelemetry 和 Python。

Distributed Tracing

Implement distributed tracing with Jaeger and Tempo for request flow visibility across microservices.

Purpose

Track requests across distributed systems to understand latency, dependencies, and failure points.

When to Use

  • Debug latency issues
  • Understand service dependencies
  • Identify bottlenecks
  • Trace error propagation
  • Analyze request paths

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Best Practices

  1. Sample appropriately (1-10% in production)
  2. Add meaningful tags (user_id, request_id)
  3. Propagate context across all service boundaries
  4. Log exceptions in spans
  5. Use consistent naming for operations
  6. Monitor tracing overhead (<1% CPU impact)
  7. Set up alerts for trace errors
  8. Implement distributed context (baggage)
  9. Use span events for important milestones
  10. Document instrumentation standards

Integration with Logging

Correlated Logs

python
import loggingfrom opentelemetry import trace
logger = logging.getLogger(__name__)
def process_request():    span = trace.get_current_span()    trace_id = span.get_span_context().trace_id
    logger.info(        "Processing request",        extra={"trace_id": format(trace_id, '032x')}    )

Troubleshooting

No traces appearing:

  • Check collector endpoint
  • Verify network connectivity
  • Check sampling configuration
  • Review application logs

High latency overhead:

  • Reduce sampling rate
  • Use batch span processor
  • Check exporter configuration

Related Skills

  • prometheus-configuration - For metrics
  • grafana-dashboards - For visualization
  • slo-implementation - For latency SLOs

来源与署名

来源:wshobson/agents位于plugins/observability-monitoring/skills/distributed-tracing提交46891e7

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架