Python Resilience

作者 wshobson46891e7e60da無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Python resilience patterns including automatic retries, exponential backoff, timeouts, and fault-tolerant decorators. Use when adding retry logic, implementing timeouts, building fault-tolerant services, or handling transient failures.

AI 產生的概覽

指導 Python 開發者為不可靠呼叫加入重試、指數退避、抖動、逾時與容錯裝飾器。

功能
此技能提供用來處理暫時性故障、網路問題與服務中斷的 Python 容錯模式。它說明暫時性錯誤與永久性錯誤的差異、指數退避、抖動與有界重試,並提供使用 tenacity 函式庫搭配 httpx 的完整範例,包括依例外、依 HTTP 狀態碼或兩者同時觸發重試。它也列出最佳實務,例如限制總時長、記錄每次重試,以及為所有呼叫設定逾時。
適用情境
適用於為外部服務呼叫加入重試邏輯、為網路操作實作逾時,或建置容錯服務與微服務。也適合處理速率限制、背壓與斷路器設計。
執行需求
僅為說明性內容,未附帶指令碼。範例假定使用 Python 以及 tenacity 與 httpx 套件,並需要網路存取以執行所示的 HTTP 呼叫。隨附的 references/details.md 檔案包含進階模式。

Python Resilience Patterns

Build fault-tolerant Python applications that gracefully handle transient failures, network issues, and service outages. Resilience patterns keep systems running when dependencies are unreliable.

When to Use This Skill

  • Adding retry logic to external service calls
  • Implementing timeouts for network operations
  • Building fault-tolerant microservices
  • Handling rate limiting and backpressure
  • Creating infrastructure decorators
  • Designing circuit breakers

Core Concepts

1. Transient vs Permanent Failures

Retry transient errors (network timeouts, temporary service issues). Don't retry permanent errors (invalid credentials, bad requests).

2. Exponential Backoff

Increase wait time between retries to avoid overwhelming recovering services.

3. Jitter

Add randomness to backoff to prevent thundering herd when many clients retry simultaneously.

4. Bounded Retries

Cap both attempt count and total duration to prevent infinite retry loops.

Quick Start

python
from tenacity import retry, stop_after_attempt, wait_exponential_jitter
@retry(    stop=stop_after_attempt(3),    wait=wait_exponential_jitter(initial=1, max=10),)def call_external_service(request: dict) -> dict:    return httpx.post("https://api.example.com", json=request).json()

Fundamental Patterns

Pattern 1: Basic Retry with Tenacity

Use the tenacity library for production-grade retry logic. For simpler cases, consider built-in retry functionality or a lightweight custom implementation.

python
from tenacity import (    retry,    stop_after_attempt,    stop_after_delay,    wait_exponential_jitter,    retry_if_exception_type,)
TRANSIENT_ERRORS = (ConnectionError, TimeoutError, OSError)
@retry(    retry=retry_if_exception_type(TRANSIENT_ERRORS),    stop=stop_after_attempt(5) | stop_after_delay(60),    wait=wait_exponential_jitter(initial=1, max=30),)def fetch_data(url: str) -> dict:    """Fetch data with automatic retry on transient failures."""    response = httpx.get(url, timeout=30)    response.raise_for_status()    return response.json()

Pattern 2: Retry Only Appropriate Errors

Whitelist specific transient exceptions. Never retry:

  • ValueError, TypeError - These are bugs, not transient issues
  • AuthenticationError - Invalid credentials won't become valid
  • HTTP 4xx errors (except 429) - Client errors are permanent
python
from tenacity import retry, retry_if_exception_typeimport httpx
# Define what's retryableRETRYABLE_EXCEPTIONS = (    ConnectionError,    TimeoutError,    httpx.ConnectTimeout,    httpx.ReadTimeout,)
@retry(    retry=retry_if_exception_type(RETRYABLE_EXCEPTIONS),    stop=stop_after_attempt(3),    wait=wait_exponential_jitter(initial=1, max=10),)def resilient_api_call(endpoint: str) -> dict:    """Make API call with retry on network issues."""    return httpx.get(endpoint, timeout=10).json()

Pattern 3: HTTP Status Code Retries

Retry specific HTTP status codes that indicate transient issues.

python
from tenacity import retry, retry_if_result, stop_after_attemptimport httpx
RETRY_STATUS_CODES = {429, 502, 503, 504}
def should_retry_response(response: httpx.Response) -> bool:    """Check if response indicates a retryable error."""    return response.status_code in RETRY_STATUS_CODES
@retry(    retry=retry_if_result(should_retry_response),    stop=stop_after_attempt(3),    wait=wait_exponential_jitter(initial=1, max=10),)def http_request(method: str, url: str, **kwargs) -> httpx.Response:    """Make HTTP request with retry on transient status codes."""    return httpx.request(method, url, timeout=30, **kwargs)

Pattern 4: Combined Exception and Status Retry

Handle both network exceptions and HTTP status codes.

python
from tenacity import (    retry,    retry_if_exception_type,    retry_if_result,    stop_after_attempt,    wait_exponential_jitter,    before_sleep_log,)import loggingimport httpx
logger = logging.getLogger(__name__)
TRANSIENT_EXCEPTIONS = (    ConnectionError,    TimeoutError,    httpx.ConnectError,    httpx.ReadTimeout,)RETRY_STATUS_CODES = {429, 500, 502, 503, 504}
def is_retryable_response(response: httpx.Response) -> bool:    return response.status_code in RETRY_STATUS_CODES
@retry(    retry=(        retry_if_exception_type(TRANSIENT_EXCEPTIONS) |        retry_if_result(is_retryable_response)    ),    stop=stop_after_attempt(5),    wait=wait_exponential_jitter(initial=1, max=30),    before_sleep=before_sleep_log(logger, logging.WARNING),)def robust_http_call(    method: str,    url: str,    **kwargs,) -> httpx.Response:    """HTTP call with comprehensive retry handling."""    return httpx.request(method, url, timeout=30, **kwargs)

Detailed worked examples and patterns

Detailed sections (starting with ## Advanced Patterns) live in references/details.md. Read that file when the navigation summary above is insufficient.

Best Practices Summary

  1. Retry only transient errors - Don't retry bugs or authentication failures
  2. Use exponential backoff - Give services time to recover
  3. Add jitter - Prevent thundering herd from synchronized retries
  4. Cap total duration - stop_after_attempt(5) | stop_after_delay(60)
  5. Log every retry - Silent retries hide systemic problems
  6. Use decorators - Keep retry logic separate from business logic
  7. Inject dependencies - Make infrastructure testable
  8. Set timeouts everywhere - Every network call needs a timeout
  9. Fail gracefully - Return cached/default values for non-critical paths
  10. Monitor retry rates - High retry rates indicate underlying issues

來源與署名

來源:wshobson/agents位於plugins/python-development/skills/python-resilience提交46891e7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架