Infrastructure

by grafana1ccacf29049fApache-2.0279 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — `k8s-monitoring` Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy `prometheus.exporter.unix` for Linux hosts, cAdvisor + Docker discovery for containers, and CloudWatch / Azure Monitor / Google Cloud Monitoring datasource setup. Use when onboarding a new cluster or VM fleet to Grafana Cloud, picking the right Helm values for K8s scraping, wiring kube-state-metrics + node-exporter + cAdvisor, alerting on `PodCrashLooping` / node memory / PVC capacity, or pulling AWS / Azure / GCP cloud metrics — even when the user says "monitor my cluster", "send K8s metrics to Grafana", "scrape EC2 metrics", "cluster pod logs", or "install the monitoring helm chart" without naming `k8s-monitoring` or Alloy.

Instructions onlyDevOps & Cloud
AI-generated overview

Onboards Kubernetes clusters, Linux hosts, containers, and cloud metrics into Grafana Cloud using Helm and Alloy.

What it does
This skill provides step-by-step instructions for shipping Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud. It covers installing the grafana/k8s-monitoring Helm chart, configuring Alloy's prometheus.exporter.unix for Linux hosts, and provisioning CloudWatch, Azure Monitor, or Google Cloud Monitoring datasources. It also includes verification commands and troubleshooting guidance for common issues.
When to use it
Use this when onboarding a new Kubernetes cluster or VM fleet to Grafana Cloud, selecting Helm values for K8s scraping, or wiring kube-state-metrics, node-exporter, and cAdvisor. It is also suited for setting up alerting on PodCrashLooping, node memory, or PVC capacity, and for pulling AWS, Azure, or GCP cloud metrics.
Requirements
Requires a Grafana Cloud stack with Prometheus, Loki, and Tempo endpoints plus an API key with metrics:write, logs:write, and traces:write permissions. For Kubernetes, needs a cluster with helm 3.x and kubectl context. For hosts or Docker, requires Alloy installed on the node. Ships no scripts; instructions only.

Grafana Cloud Infrastructure Monitoring

Docs: https://grafana.com/docs/grafana-cloud/monitor-infrastructure/

K8s + host + container + cloud-provider telemetry, mostly via the grafana/k8s-monitoring Helm chart or Alloy.

Prerequisites

  • Grafana Cloud stack with Prometheus / Loki / Tempo endpoints + API key (metrics:write, logs:write, traces:write)
  • For Kubernetes: a cluster + helm 3.x + kubectl context pointing at it
  • For hosts / Docker: Alloy installed on the node

Common Workflows

1. Onboard a Kubernetes cluster (k8s-monitoring chart)

bash
# 1. Create the namespace + secretkubectl create namespace monitoringkubectl create secret generic grafana-cloud-secret \  -n monitoring --from-literal=api-key=<your-api-key>
# 2. Install — values.yaml in references/k8s-monitoring-values.mdhelm repo add grafana https://grafana.github.io/helm-charts && helm repo updatehelm install k8s-monitoring grafana/k8s-monitoring \  --version 4.1.4 -n monitoring -f values.yaml
# 3. Verify every pod is Runningkubectl get pods -n monitoring# Expect alloy-*, kube-state-metrics-*, node-exporter-*, etc. all Ready.
# 4. Verify no error logs in the metrics/logs/traces Alloyskubectl -n monitoring logs deploy/k8s-monitoring-alloy-metrics --tail=50 | grep -iE 'error|level=err' || echo "clean"
# 5. Verify telemetry landed in Grafana Cloud#    PromQL on the metrics datasource (should be > 0):#      sum(up{cluster="production-us-east"})#    LogQL on Loki:#      sum(count_over_time({cluster="production-us-east"}[5m]))

Full values.yaml, key PromQL, dashboard IDs (15520, 1860, 14282…), and alert rules: references/k8s-monitoring-values.md [blocked].

2. Monitor a Linux host

alloy
# 1. /etc/alloy/config.alloy — see references/clouds-and-hosts.md for the full blockprometheus.exporter.unix "host"  { rootfs_path = "/" }prometheus.scrape         "node" { targets = prometheus.exporter.unix.host.targets                                   forward_to = [prometheus.remote_write.cloud.receiver] }
bash
# 2. Reload Alloy and verify the unix exporter is upsystemctl reload alloycurl -s http://localhost:12345/api/v0/web/components | jq '.[] | select(.id|contains("prometheus.exporter.unix"))'
# 3. Verify in Grafana Cloud — open the "Node Exporter Full" dashboard (ID 1860)#    and pick your host from the `instance` dropdown.

3. Pull AWS / Azure / GCP metrics

Provision the datasource (full YAML in references/clouds-and-hosts.md [blocked]), then:

bash
# 1. After provisioning, restart Grafana to pick up the file# 2. Verify the datasource — Grafana → Connections → Data sources → "Test"#    Expect "Successfully queried the CloudWatch metrics API" (or equivalent).# 3. Confirm a query — Explore → datasource → metric e.g.#    CloudWatch namespace AWS/EC2 metric CPUUtilization, last 1h.

Troubleshooting

  • chart installed but no metrics in Cloud → check the grafana-cloud-secret api-key value; check Alloy logs for 401
  • kube-state-metrics pod Pending → likely RBAC; reapply the chart's CRDs/CRBs
  • Node-exporter pod CrashLoopBackOff → typically hostNetwork: true collision with the host's :9100; change the port
  • CloudWatch "Access denied" → IAM role missing cloudwatch:GetMetricData, cloudwatch:ListMetrics

Resources

Source and attribution

Source:grafana/skillsinskills/grafana-cloud/infrastructureat commit1ccacf2

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from grafana/skills

React 19 Plugin Migration

grafana

Guides migration of a Grafana plugin to React 19 compatibility through ordered build, dependency and source-code steps.

Software Development279updated today

Plugin Bundle Size

grafana

Guides optimisation of Grafana app plugin bundle size using React.lazy, Suspense and webpack code splitting.

Software Development279updated today

Grafana Scenes

grafana

Builds Grafana plugin pages with the @grafana/scenes framework, covering scenes, panels, variables and drilldowns.

Software Development279updated today

Check Npm

grafana

Read-only audit of npm, yarn, or pnpm configuration for supply-chain hardening in a JS/TS repository.

Security279updated today

Mimir

grafana

Guides standing up and operating Grafana Mimir for scalable, multi-tenant, long-term Prometheus and OTLP metrics storage.

DevOps & Cloud279updated today

K6 Trend Analysis

grafana

Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for performance regression", "should I tighten thresholds", "are my tests degrading", "show me trends for test X", "analyze my k6 test runs", or "is my test getting slower". Also trigger when a user asks to check all tests in a project -- run this skill once per test and synthesize.

Awaiting classification279updated today