Infrastructure Monitoring

aj-geddes/useful-ai-prompts/skills/infrastructure-monitoring

作者 aj-geddes3f5182cfd739無授權條款355 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫7 個月前更新

Set up comprehensive infrastructure monitoring with Prometheus, Grafana, and alerting systems for metrics, health checks, and performance tracking.

包含腳本DevOps & Cloud
AI 產生的概覽

使用 Prometheus、Grafana 與告警系統建置基礎架構監控,用於指標、健康檢查與效能追蹤。

功能
提供部署監控堆疊的設定指引與參考資料:Prometheus 抓取與規則設定、告警規則、Alertmanager 設定、Grafana 儀表板以及部署步驟。同時附上一個健康檢查 shell 指令碼和一個儀表板設定範本。最終產出可用於追蹤系統健康與資源使用情況的監控與告警設定。
適用情境
適用於為伺服器與服務建置或擴充監控的情境,包括即時效能追蹤、容量規劃、事件偵測與告警,以及資源使用率分析。
執行需求
需要 Prometheus、Grafana 與 Alertmanager(或相容部署);執行隨附健康檢查指令碼需要 shell 環境;需要設定檔以及存取被監控目標的網路。

Infrastructure Monitoring

Table of Contents

Overview

Implement comprehensive infrastructure monitoring to track system health, performance metrics, and resource utilization with alerting and visualization across your entire stack.

When to Use

  • Real-time performance monitoring
  • Capacity planning and trends
  • Incident detection and alerting
  • Service health tracking
  • Resource utilization analysis
  • Performance troubleshooting
  • Compliance and audit trails
  • Historical data analysis

Quick Start

Minimal working example:

yaml
# prometheus.ymlglobal:  scrape_interval: 15s  evaluation_interval: 15s  external_labels:    monitor: "infrastructure-monitor"    environment: "production"
# Alertmanager configurationalerting:  alertmanagers:    - static_configs:        - targets:            - localhost:9093
# Rule filesrule_files:  - "alerts.yml"  - "rules.yml"
scrape_configs:  # Prometheus itself  - job_name: "prometheus"    static_configs:      - targets: ["localhost:9090"]// ... (see reference guides for full implementation)

Reference Guides

Detailed implementations in the references/ directory:

GuideContents
Prometheus Configuration [blocked]Prometheus Configuration
Alert Rules [blocked]Alert Rules
Alertmanager Configuration [blocked]Alertmanager Configuration
Grafana Dashboard [blocked]Grafana Dashboard
Monitoring Deployment [blocked]Monitoring Deployment

Best Practices

✅ DO

  • Follow established patterns and conventions
  • Write clean, maintainable code
  • Add appropriate documentation
  • Test thoroughly before deploying

❌ DON'T

  • Skip testing or validation
  • Ignore error handling
  • Hard-code configuration values

來源與署名

來源:aj-geddes/useful-ai-prompts位於skills/infrastructure-monitoring提交3f5182c

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架