Cost Aware Llm Pipeline

affaan-m/ECC/skills/cost-aware-llm-pipeline

作者 affaan-mef648e01899ba3e8dc6371642deaaf64b4477775無授權條款275K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫4 天前更新

Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching. Use when LLM spend needs to come down, or when routing tasks across model tiers and budgets.

僅含說明AI & Agents
AI 產生的概覽

透過模型路由、預算追蹤、重試邏輯與提示快取來降低 LLM API 成本的模式。

功能
提供建構成本感知 LLM API 管線的參考模式與程式碼片段:依任務複雜度在便宜與昂貴模型之間路由、以不可變記錄追蹤累計花費、僅對暫時性錯誤重試,以及快取較長的系統提示。也包含模型定價參考表與最佳實務及反模式清單。產出的是指引與範例程式碼,而非可直接執行的工具。
適用情境
適用於需要降低 LLM API 支出、需要跨模型層級路由任務,或批次處理需要預算護欄的情境。適合呼叫 Claude、OpenAI 或類似 API 的應用,以及需要智慧路由的多模型架構。
執行需求
此技能不附帶指令碼,僅為說明與程式碼範例。範例涉及 Python、anthropic 用戶端程式庫及其錯誤類型,並假定可存取 LLM API。

Cost-Aware LLM Pipeline

Patterns for controlling LLM API costs while maintaining quality. Combines model routing, budget tracking, retry logic, and prompt caching into a composable pipeline.

When to Activate

  • Building applications that call LLM APIs (Claude, GPT, etc.)
  • Processing batches of items with varying complexity
  • Need to stay within a budget for API spend
  • Optimizing cost without sacrificing quality on complex tasks

Core Concepts

1. Model Routing by Task Complexity

Automatically select cheaper models for simple tasks, reserving expensive models for complex ones.

python
MODEL_SONNET = "claude-sonnet-5"MODEL_HAIKU = "claude-haiku-4-5-20251001"
_SONNET_TEXT_THRESHOLD = 10_000  # chars_SONNET_ITEM_THRESHOLD = 30     # items
def select_model(    text_length: int,    item_count: int,    force_model: str | None = None,) -> str:    """Select model based on task complexity."""    if force_model is not None:        return force_model    if text_length >= _SONNET_TEXT_THRESHOLD or item_count >= _SONNET_ITEM_THRESHOLD:        return MODEL_SONNET  # Complex task    return MODEL_HAIKU  # Simple task (3-4x cheaper)

2. Immutable Cost Tracking

Track cumulative spend with frozen dataclasses. Each API call returns a new tracker — never mutates state.

python
from dataclasses import dataclass
@dataclass(frozen=True, slots=True)class CostRecord:    model: str    input_tokens: int    output_tokens: int    cost_usd: float
@dataclass(frozen=True, slots=True)class CostTracker:    budget_limit: float = 1.00    records: tuple[CostRecord, ...] = ()
    def add(self, record: CostRecord) -> "CostTracker":        """Return new tracker with added record (never mutates self)."""        return CostTracker(            budget_limit=self.budget_limit,            records=(*self.records, record),        )
    @property    def total_cost(self) -> float:        return sum(r.cost_usd for r in self.records)
    @property    def over_budget(self) -> bool:        return self.total_cost > self.budget_limit

3. Narrow Retry Logic

Retry only on transient errors. Fail fast on authentication or bad request errors.

python
from anthropic import (    APIConnectionError,    InternalServerError,    RateLimitError,)
_RETRYABLE_ERRORS = (APIConnectionError, RateLimitError, InternalServerError)_MAX_RETRIES = 3
def call_with_retry(func, *, max_retries: int = _MAX_RETRIES):    """Retry only on transient errors, fail fast on others."""    for attempt in range(max_retries):        try:            return func()        except _RETRYABLE_ERRORS:            if attempt == max_retries - 1:                raise            time.sleep(2 ** attempt)  # Exponential backoff    # AuthenticationError, BadRequestError etc. → raise immediately

4. Prompt Caching

Cache long system prompts to avoid resending them on every request.

python
messages = [    {        "role": "user",        "content": [            {                "type": "text",                "text": system_prompt,                "cache_control": {"type": "ephemeral"},  # Cache this            },            {                "type": "text",                "text": user_input,  # Variable part            },        ],    }]

Composition

Combine all four techniques in a single pipeline function:

python
def process(text: str, config: Config, tracker: CostTracker) -> tuple[Result, CostTracker]:    # 1. Route model    model = select_model(len(text), estimated_items, config.force_model)
    # 2. Check budget    if tracker.over_budget:        raise BudgetExceededError(tracker.total_cost, tracker.budget_limit)
    # 3. Call with retry + caching    response = call_with_retry(lambda: client.messages.create(        model=model,        messages=build_cached_messages(system_prompt, text),    ))
    # 4. Track cost (immutable)    record = CostRecord(model=model, input_tokens=..., output_tokens=..., cost_usd=...)    tracker = tracker.add(record)
    return parse_result(response), tracker

Pricing Reference (2026)

ModelInput ($/1M tokens)Output ($/1M tokens)Relative Cost
Haiku 3.5 (legacy)$0.80$4.000.8x
Haiku 4.5$1.00$5.001x
Sonnet 5$2.00$10.002x
Sonnet 4.6$3.00$15.003x
Opus 4.8$5.00$25.005x
Fable 5 / Mythos 5$10.00$50.0010x
Opus 4.0 / 4.1 (legacy)$15.00$75.0015x

Best Practices

  • Start with the cheapest model and only route to expensive models when complexity thresholds are met
  • Set explicit budget limits before processing batches — fail early rather than overspend
  • Log model selection decisions so you can tune thresholds based on real data
  • Use prompt caching for system prompts over 1024 tokens — saves both cost and latency
  • Never retry on authentication or validation errors — only transient failures (network, rate limit, server error)

Anti-Patterns to Avoid

  • Using the most expensive model for all requests regardless of complexity
  • Retrying on all errors (wastes budget on permanent failures)
  • Mutating cost tracking state (makes debugging and auditing difficult)
  • Hardcoding model names throughout the codebase (use constants or config)
  • Ignoring prompt caching for repetitive system prompts

When to Use

  • Any application calling Claude, OpenAI, or similar LLM APIs
  • Batch processing pipelines where cost adds up quickly
  • Multi-model architectures that need intelligent routing
  • Production systems that need budget guardrails

來源與署名

來源:affaan-m/ECC位於skills/cost-aware-llm-pipeline提交ef648e0

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架