Guides operation of long-running, cloud-hosted agent workloads through observability, security boundaries and lifecycle management.
- What it does
- This skill provides operational guidance for agent systems that run continuously or in the cloud rather than in a single CLI session. It covers runtime lifecycle (start, pause, stop, restart), observability (logs, metrics, traces), security controls (scoping, permissions, kill switches) and change management (release, rollback, audit). It also lists baseline controls, metrics to track, an incident-handling pattern and deployment integrations.
- When to use it
- Use it when operating cloud-hosted or continuously running agent systems that need control beyond a single session. It suits teams setting up lifecycle, observability, security and change-management practices for such workloads, or responding to failure spikes.
- Requirements
- Instructions only; no scripts are included. It references working with PM2 workflows, systemd services, container orchestrators and CI/CD gates, which implies access to those environments.