Python Observability

作者 wshobson46891e7e60da无许可证收录于 2026年10月8日更新于 2026年10月8日

Python observability patterns including structured logging, metrics, and distributed tracing. Use when adding logging, implementing metrics collection, setting up tracing, or debugging production systems.

AI 生成的概览

指导为 Python 应用添加结构化日志、指标和分布式追踪,便于排查生产问题。

功能
提供模式与代码示例,用于添加基于 structlog 的结构化 JSON 日志、统一日志字段、语义化日志级别,以及跨服务传递关联 ID。还涵盖指标概念,如四大黄金信号与标签基数控制,以及分布式追踪的搭建。详细的完整示例放在单独的参考文件中。
适用场景
适用于添加日志、实现指标采集、搭建追踪、传递关联 ID 或排查生产系统问题的场景。目标是让 Python 服务具备可观测性,从而无需重新部署代码即可定位生产问题。
运行要求
需要 Python 及 structlog 库;示例中还涉及 FastAPI 和 httpx。不附带脚本,仅为说明与代码片段,另有一个可选参考文件。

Python Observability

Instrument Python applications with structured logs, metrics, and traces. When something breaks in production, you need to answer "what, where, and why" without deploying new code.

When to Use This Skill

  • Adding structured logging to applications
  • Implementing metrics collection with Prometheus
  • Setting up distributed tracing across services
  • Propagating correlation IDs through request chains
  • Debugging production issues
  • Building observability dashboards

Core Concepts

1. Structured Logging

Emit logs as JSON with consistent fields for production environments. Machine-readable logs enable powerful queries and alerts. For local development, consider human-readable formats.

2. The Four Golden Signals

Track latency, traffic, errors, and saturation for every service boundary.

3. Correlation IDs

Thread a unique ID through all logs and spans for a single request, enabling end-to-end tracing.

4. Bounded Cardinality

Keep metric label values bounded. Unbounded labels (like user IDs) explode storage costs.

Quick Start

python
import structlog
structlog.configure(    processors=[        structlog.processors.TimeStamper(fmt="iso"),        structlog.processors.JSONRenderer(),    ],)
logger = structlog.get_logger()logger.info("Request processed", user_id="123", duration_ms=45)

Fundamental Patterns

Pattern 1: Structured Logging with Structlog

Configure structlog for JSON output with consistent fields.

python
import loggingimport structlog
def configure_logging(log_level: str = "INFO") -> None:    """Configure structured logging for the application."""    structlog.configure(        processors=[            structlog.contextvars.merge_contextvars,            structlog.processors.add_log_level,            structlog.processors.TimeStamper(fmt="iso"),            structlog.processors.StackInfoRenderer(),            structlog.processors.format_exc_info,            structlog.processors.JSONRenderer(),        ],        wrapper_class=structlog.make_filtering_bound_logger(            getattr(logging, log_level.upper())        ),        context_class=dict,        logger_factory=structlog.PrintLoggerFactory(),        cache_logger_on_first_use=True,    )
# Initialize at application startupconfigure_logging("INFO")logger = structlog.get_logger()

Pattern 2: Consistent Log Fields

Every log entry should include standard fields for filtering and correlation.

python
import structlogfrom contextvars import ContextVar
# Store correlation ID in contextcorrelation_id: ContextVar[str] = ContextVar("correlation_id", default="")
logger = structlog.get_logger()
def process_request(request: Request) -> Response:    """Process request with structured logging."""    logger.info(        "Request received",        correlation_id=correlation_id.get(),        method=request.method,        path=request.path,        user_id=request.user_id,    )
    try:        result = handle_request(request)        logger.info(            "Request completed",            correlation_id=correlation_id.get(),            status_code=200,            duration_ms=elapsed,        )        return result    except Exception as e:        logger.error(            "Request failed",            correlation_id=correlation_id.get(),            error_type=type(e).__name__,            error_message=str(e),        )        raise

Pattern 3: Semantic Log Levels

Use log levels consistently across the application.

LevelPurposeExamples
DEBUGDevelopment diagnosticsVariable values, internal state
INFORequest lifecycle, operationsRequest start/end, job completion
WARNINGRecoverable anomaliesRetry attempts, fallback used
ERRORFailures needing attentionExceptions, service unavailable
python
# DEBUG: Detailed internal informationlogger.debug("Cache lookup", key=cache_key, hit=cache_hit)
# INFO: Normal operational eventslogger.info("Order created", order_id=order.id, total=order.total)
# WARNING: Abnormal but handled situationslogger.warning(    "Rate limit approaching",    current_rate=950,    limit=1000,    reset_seconds=30,)
# ERROR: Failures requiring investigationlogger.error(    "Payment processing failed",    order_id=order.id,    error=str(e),    payment_provider="stripe",)

Never log expected behavior at ERROR. A user entering a wrong password is INFO, not ERROR.

Pattern 4: Correlation ID Propagation

Generate a unique ID at ingress and thread it through all operations.

python
from contextvars import ContextVarimport uuidimport structlog
correlation_id: ContextVar[str] = ContextVar("correlation_id", default="")
def set_correlation_id(cid: str | None = None) -> str:    """Set correlation ID for current context."""    cid = cid or str(uuid.uuid4())    correlation_id.set(cid)    structlog.contextvars.bind_contextvars(correlation_id=cid)    return cid
# FastAPI middleware examplefrom fastapi import Request
async def correlation_middleware(request: Request, call_next):    """Middleware to set and propagate correlation ID."""    # Use incoming header or generate new    cid = request.headers.get("X-Correlation-ID") or str(uuid.uuid4())    set_correlation_id(cid)
    response = await call_next(request)    response.headers["X-Correlation-ID"] = cid    return response

Propagate to outbound requests:

python
import httpx
async def call_downstream_service(endpoint: str, data: dict) -> dict:    """Call downstream service with correlation ID."""    async with httpx.AsyncClient() as client:        response = await client.post(            endpoint,            json=data,            headers={"X-Correlation-ID": correlation_id.get()},        )        return response.json()

Detailed worked examples and patterns

Detailed sections (starting with ## Advanced Patterns) live in references/details.md. Read that file when the navigation summary above is insufficient.

Best Practices Summary

  1. Use structured logging - JSON logs with consistent fields
  2. Propagate correlation IDs - Thread through all requests and logs
  3. Track the four golden signals - Latency, traffic, errors, saturation
  4. Bound label cardinality - Never use unbounded values as metric labels
  5. Log at appropriate levels - Don't cry wolf with ERROR
  6. Include context - User ID, request ID, operation name in logs
  7. Use context managers - Consistent timing and error handling
  8. Separate concerns - Observability code shouldn't pollute business logic
  9. Test your observability - Verify logs and metrics in integration tests
  10. Set up alerts - Metrics are useless without alerting

来源与署名

来源:wshobson/agents位于plugins/python-development/skills/python-observability提交46891e7

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架