Enterprise Agent Ops

by affaan-mef648e01899bNo license275K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 days ago

オブザーバビリティ、セキュリティ境界、およびライフサイクル管理を備えた長寿命エージェントワークロードを運用します。

Instructions onlyDevOps & CloudAI & Agents
AI-generated overview

Operational guidance for running long-lived, cloud-hosted agent workloads with observability, safety controls and lifecycle management.

What it does
This skill provides operational guidance for continuously running or cloud-hosted agent systems. It covers four domains: runtime lifecycle (start, pause, stop, restart), observability (logs, metrics, traces), safety controls (scopes, permissions, kill switches) and change management (rollout, rollback, audit). It also lists baseline controls, metrics to track, an incident response pattern and deployment integrations.
When to use it
Use it when operating agent systems that run beyond a single CLI session and need operational controls. It fits teams managing deployment, monitoring, incident response and change management for hosted agents.
Requirements
No scripts are included; it is instructions only. It references pairing with PM2 workflows, systemd services, container orchestrators and CI/CD gates, but requires no specific tools, packages or credentials to read.

Enterprise Agent Ops

Use this skill for cloud-hosted or continuously running agent systems that need operational controls beyond single CLI sessions.

Operational Domains

  1. runtime lifecycle (start, pause, stop, restart)
  2. observability (logs, metrics, traces)
  3. safety controls (scopes, permissions, kill switches)
  4. change management (rollout, rollback, audit)

Baseline Controls

  • immutable deployment artifacts
  • least-privilege credentials
  • environment-level secret injection
  • hard timeout and retry budgets
  • audit log for high-risk actions

Metrics to Track

  • success rate
  • mean retries per task
  • time to recovery
  • cost per successful task
  • failure class distribution

Incident Pattern

When failure spikes:

  1. freeze new rollout
  2. capture representative traces
  3. isolate failing route
  4. patch with smallest safe change
  5. run regression + security checks
  6. resume gradually

Deployment Integrations

This skill pairs with:

  • PM2 workflows
  • systemd services
  • container orchestrators
  • CI/CD gates

Source and attribution

Source:affaan-m/eccindocs/ja-JP/skills/enterprise-agent-opsat commitef648e0

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal