Infrastructure Monitoring

aj-geddes/useful-ai-prompts/skills/infrastructure-monitoring

作者 aj-geddes3f5182cfd739无许可证355 个星标收录于 2026年10月8日更新于 2026年10月8日仓库7个月前更新

Set up comprehensive infrastructure monitoring with Prometheus, Grafana, and alerting systems for metrics, health checks, and performance tracking.

包含脚本DevOps & Cloud
AI 生成的概览

使用 Prometheus、Grafana 和告警系统搭建基础设施监控,用于指标、健康检查与性能跟踪。

功能
提供部署监控栈的配置指导与参考资料:Prometheus 抓取与规则配置、告警规则、Alertmanager 配置、Grafana 仪表板以及部署步骤。同时附带一个健康检查 shell 脚本和一个仪表板配置模板。最终产出可用于跟踪系统健康与资源使用情况的监控与告警配置。
适用场景
适用于为服务器和服务搭建或扩展监控的场景,包括实时性能跟踪、容量规划、事件检测与告警,以及资源使用率分析。
运行要求
需要 Prometheus、Grafana 和 Alertmanager(或兼容部署);运行随附健康检查脚本需要 shell 环境;需要配置文件以及访问被监控目标的网络。

Infrastructure Monitoring

Table of Contents

Overview

Implement comprehensive infrastructure monitoring to track system health, performance metrics, and resource utilization with alerting and visualization across your entire stack.

When to Use

  • Real-time performance monitoring
  • Capacity planning and trends
  • Incident detection and alerting
  • Service health tracking
  • Resource utilization analysis
  • Performance troubleshooting
  • Compliance and audit trails
  • Historical data analysis

Quick Start

Minimal working example:

yaml
# prometheus.ymlglobal:  scrape_interval: 15s  evaluation_interval: 15s  external_labels:    monitor: "infrastructure-monitor"    environment: "production"
# Alertmanager configurationalerting:  alertmanagers:    - static_configs:        - targets:            - localhost:9093
# Rule filesrule_files:  - "alerts.yml"  - "rules.yml"
scrape_configs:  # Prometheus itself  - job_name: "prometheus"    static_configs:      - targets: ["localhost:9090"]// ... (see reference guides for full implementation)

Reference Guides

Detailed implementations in the references/ directory:

GuideContents
Prometheus Configuration [blocked]Prometheus Configuration
Alert Rules [blocked]Alert Rules
Alertmanager Configuration [blocked]Alertmanager Configuration
Grafana Dashboard [blocked]Grafana Dashboard
Monitoring Deployment [blocked]Monitoring Deployment

Best Practices

✅ DO

  • Follow established patterns and conventions
  • Write clean, maintainable code
  • Add appropriate documentation
  • Test thoroughly before deploying

❌ DON'T

  • Skip testing or validation
  • Ignore error handling
  • Hard-code configuration values

来源与署名

来源:aj-geddes/useful-ai-prompts位于skills/infrastructure-monitoring提交3f5182c

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架