Vtex Io Observability And Ops

vtex/skills/skills/vtex-io-observability-and-ops

作者 vtex799a5a24c3da679ceb40dfe9e690023b05d108fc无许可证收录于 2026年10月9日更新于 2026年10月9日

Apply when making VTEX IO services easier to observe, troubleshoot, and operate in production. Covers metrics, structured logging, failure visibility, rate-limit awareness, and production readiness checks for backend apps. Use for integration monitoring, error diagnosis, or improving the operational quality of VTEX IO services before or after release.

仅含说明DevOps & Cloud
AI 生成的概览

指导 VTEX IO 后端服务的可观测性与运维就绪:指标、结构化日志与故障可见性。

功能
提供决策规则和硬性约束,使 VTEX IO 服务在生产环境中可观测、可运维。内容涵盖在重要集成调用上发出客户端级指标、使用 ctx.vtex.logger 并携带结构化上下文而非 console.log、避免静默吞掉故障、记录日志前脱敏敏感数据,以及针对易触发限流的集成采用超时、退避重试与缓存。还包含生产就绪检查清单和常见故障模式。
适用场景
适用于为客户端调用或流程添加指标、改进路由/worker/集成的日志、向运维与支持团队暴露故障、评审服务是否可上线,或监控对限流敏感的集成。不适用于应用策略声明、信任边界建模、前端分析或单独的路由契约设计。
运行要求
无需脚本或特殊工具,仅为说明性内容。假定处于 VTEX IO 后端应用环境,可访问 ctx.vtex.logger 与 Node HTTP 客户端,并引用 VTEX 开发者文档。

Observability & Operational Readiness

When this skill applies

Use this skill when a VTEX IO service needs better production visibility, troubleshooting behavior, or operational safety.

  • Adding metrics to important client calls or flows
  • Improving logs for routes, workers, or integrations
  • Surfacing failures clearly for operations and support
  • Reviewing whether a service is ready for production
  • Monitoring rate-limit-sensitive integrations

Do not use this skill for:

  • app policy declaration
  • trust-boundary modeling
  • frontend analytics or browser monitoring
  • route contract design by itself

Decision rules

  • Log enough structured context to debug failures, but do not log secrets or sensitive payloads.
  • Use ctx.vtex.logger with appropriate log levels such as info, warn, and error instead of console.log, so logs are properly collected and searchable in the VTEX logging stack.
  • Treat ctx.vtex.logger as the native platform logging mechanism. If a partner needs to forward logs to its own logging system, prefer doing that through a dedicated integration app or client instead of replacing the VTEX logger pattern inside every service.
  • Use client-level metrics on important downstream calls so integration behavior is visible below the handler layer.
  • Choose metric names that reflect the integration and operation, such as partner-get-order or partner-sync-catalog, so counts, latency, and error rates can be tracked over time.
  • Make failures observable at the point where they happen. Do not swallow errors silently in routes, events, or workers.
  • For rate-limit-sensitive APIs, combine short timeouts, backoff-aware retries, and caching of frequent reads to reduce burst pressure and avoid hitting hard limits.
  • Review whether expensive or fragile flows expose enough operational signals before releasing them.

Hard constraints

Constraint: Important failures must be visible in logs, metrics, or durable state

Routes, event handlers, and workers MUST not hide important failures from operators.

Why this matters

If failures disappear silently, the service becomes impossible to diagnose under real traffic and retries.

Detection

If an error is caught and ignored without logging, metric emission, or explicit failure state, STOP and surface the failure.

Correct

typescript
try {  await ctx.clients.partnerApi.sendOrder(orderId)} catch (error) {  ctx.vtex.logger.error({    message: 'Failed to send order to partner',    orderId,    account: ctx.vtex.account,    routeId: ctx.vtex.route?.id,  })  throw error}

Wrong

typescript
try {  await ctx.clients.partnerApi.sendOrder(orderId)} catch (_) {  return}

Constraint: Metrics should be attached to important integration calls

Client calls that are operationally important SHOULD include metric so request behavior can be tracked consistently.

Why this matters

Without metrics, integration failures and latency patterns are much harder to isolate from generic route behavior.

Detection

If a key downstream integration call has no metric and operations depend on it, STOP and add a meaningful metric name.

Correct

typescript
return this.http.get(`/orders/${id}`, {  metric: 'partner-get-order',})

Wrong

typescript
return this.http.get(`/orders/${id}`)

Constraint: Logs must stay useful without leaking sensitive data

Logs MUST contain enough context to debug production behavior, but MUST NOT include secrets, tokens, or unnecessarily sensitive payloads.

Why this matters

Operational logs are only valuable if they are safe to retain and inspect. Sensitive logging creates security risk while still failing to guarantee useful diagnosis.

Detection

If a log line includes tokens, auth headers, raw personal payloads, or entire downstream responses, STOP and sanitize the log.

Correct

typescript
ctx.vtex.logger.info({  message: 'Partner sync started',  orderId,  account: ctx.vtex.account,})

Wrong

typescript
ctx.vtex.logger.info({  message: 'Partner sync started',  body: ctx.request.body,  auth: ctx.request.header.authorization,})

Preferred pattern

Operationally healthy VTEX IO services should:

  • emit metrics for important client calls so counts, latency, and error rates are visible
  • log failures with enough structured context such as domain IDs, account, and routeId
  • avoid silent error swallowing
  • sanitize sensitive data before logging
  • review retries, caching, and throughput with rate-limit behavior in mind

Use observability to shorten diagnosis time, not just to create more logs.

Common failure modes

  • Catching and ignoring errors in async flows.
  • Logging too little context to diagnose production incidents.
  • Logging too much sensitive data.
  • Omitting metrics from important integration calls.
  • Treating rate-limit failures as isolated bugs instead of operational signals.

Review checklist

  • Are important failures visible to operators?
  • Do key integrations emit useful metrics?
  • Are logs structured and safe?
  • Are retries, caching, and rate-limit behavior considered together?
  • Would someone on call be able to diagnose this flow from the available signals?

Reference

来源与署名

来源:vtex/skills位于skills/vtex-io-observability-and-ops提交799a5a2

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架