Infrastructure Monitoring

by aj-geddes3f5182cfd739No license355 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 7 months ago

Set up comprehensive infrastructure monitoring with Prometheus, Grafana, and alerting systems for metrics, health checks, and performance tracking.

Includes scriptsDevOps & Cloud
AI-generated overview

Sets up infrastructure monitoring with Prometheus, Grafana, and alerting for metrics, health checks, and performance tracking.

What it does
Provides configuration guidance and reference material for deploying a monitoring stack: Prometheus scrape and rule configuration, alert rules, Alertmanager configuration, Grafana dashboards, and deployment steps. It also ships a health-check shell script and a dashboard configuration template. The result is a working monitoring and alerting setup for tracking system health and resource usage.
When to use it
Use when standing up or extending monitoring for servers and services, including real-time performance tracking, capacity planning, incident detection and alerting, and resource utilization analysis.
Requirements
Prometheus, Grafana, and Alertmanager (or compatible deployments); shell environment to run the bundled health-check script; configuration files and network access to the monitored targets.

Infrastructure Monitoring

Table of Contents

Overview

Implement comprehensive infrastructure monitoring to track system health, performance metrics, and resource utilization with alerting and visualization across your entire stack.

When to Use

  • Real-time performance monitoring
  • Capacity planning and trends
  • Incident detection and alerting
  • Service health tracking
  • Resource utilization analysis
  • Performance troubleshooting
  • Compliance and audit trails
  • Historical data analysis

Quick Start

Minimal working example:

yaml
# prometheus.ymlglobal:  scrape_interval: 15s  evaluation_interval: 15s  external_labels:    monitor: "infrastructure-monitor"    environment: "production"
# Alertmanager configurationalerting:  alertmanagers:    - static_configs:        - targets:            - localhost:9093
# Rule filesrule_files:  - "alerts.yml"  - "rules.yml"
scrape_configs:  # Prometheus itself  - job_name: "prometheus"    static_configs:      - targets: ["localhost:9090"]// ... (see reference guides for full implementation)

Reference Guides

Detailed implementations in the references/ directory:

GuideContents
Prometheus Configuration [blocked]Prometheus Configuration
Alert Rules [blocked]Alert Rules
Alertmanager Configuration [blocked]Alertmanager Configuration
Grafana Dashboard [blocked]Grafana Dashboard
Monitoring Deployment [blocked]Monitoring Deployment

Best Practices

✅ DO

  • Follow established patterns and conventions
  • Write clean, maintainable code
  • Add appropriate documentation
  • Test thoroughly before deploying

❌ DON'T

  • Skip testing or validation
  • Ignore error handling
  • Hard-code configuration values

Source and attribution

Source:aj-geddes/useful-ai-promptsinskills/infrastructure-monitoringat commit3f5182c

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal