Infrastructure

作者 grafana1ccacf29049fApache-2.0279 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — `k8s-monitoring` Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy `prometheus.exporter.unix` for Linux hosts, cAdvisor + Docker discovery for containers, and CloudWatch / Azure Monitor / Google Cloud Monitoring datasource setup. Use when onboarding a new cluster or VM fleet to Grafana Cloud, picking the right Helm values for K8s scraping, wiring kube-state-metrics + node-exporter + cAdvisor, alerting on `PodCrashLooping` / node memory / PVC capacity, or pulling AWS / Azure / GCP cloud metrics — even when the user says "monitor my cluster", "send K8s metrics to Grafana", "scrape EC2 metrics", "cluster pod logs", or "install the monitoring helm chart" without naming `k8s-monitoring` or Alloy.

僅含說明DevOps & Cloud
AI 產生的概覽

使用 Helm 和 Alloy 將 Kubernetes 叢集、Linux 主機、容器與雲端指標接入 Grafana Cloud。

功能
此技能提供將 Kubernetes、主機、容器與雲端供應商遙測資料傳送至 Grafana Cloud 的逐步說明。涵蓋安裝 grafana/k8s-monitoring Helm chart、為 Linux 主機設定 Alloy 的 prometheus.exporter.unix,以及佈建 CloudWatch、Azure Monitor 或 Google Cloud Monitoring 資料來源。也包含驗證指令與常見問題的疑難排解指引。
適用情境
適用於將新的 Kubernetes 叢集或虛擬機機群接入 Grafana Cloud、為 K8s 抓取選擇 Helm values,或配置 kube-state-metrics、node-exporter 與 cAdvisor。也適合設定 PodCrashLooping、節點記憶體或 PVC 容量告警,以及拉取 AWS、Azure 或 GCP 雲端指標。
執行需求
需要具有 Prometheus、Loki 與 Tempo 端點以及具備 metrics:write、logs:write 與 traces:write 權限的 API 金鑰的 Grafana Cloud 堆疊。對於 Kubernetes,需要叢集以及 helm 3.x 與 kubectl 內容。對於主機或 Docker,需要在節點上安裝 Alloy。不附帶指令碼;僅為說明。

Grafana Cloud Infrastructure Monitoring

Docs: https://grafana.com/docs/grafana-cloud/monitor-infrastructure/

K8s + host + container + cloud-provider telemetry, mostly via the grafana/k8s-monitoring Helm chart or Alloy.

Prerequisites

  • Grafana Cloud stack with Prometheus / Loki / Tempo endpoints + API key (metrics:write, logs:write, traces:write)
  • For Kubernetes: a cluster + helm 3.x + kubectl context pointing at it
  • For hosts / Docker: Alloy installed on the node

Common Workflows

1. Onboard a Kubernetes cluster (k8s-monitoring chart)

bash
# 1. Create the namespace + secretkubectl create namespace monitoringkubectl create secret generic grafana-cloud-secret \  -n monitoring --from-literal=api-key=<your-api-key>
# 2. Install — values.yaml in references/k8s-monitoring-values.mdhelm repo add grafana https://grafana.github.io/helm-charts && helm repo updatehelm install k8s-monitoring grafana/k8s-monitoring \  --version 4.1.4 -n monitoring -f values.yaml
# 3. Verify every pod is Runningkubectl get pods -n monitoring# Expect alloy-*, kube-state-metrics-*, node-exporter-*, etc. all Ready.
# 4. Verify no error logs in the metrics/logs/traces Alloyskubectl -n monitoring logs deploy/k8s-monitoring-alloy-metrics --tail=50 | grep -iE 'error|level=err' || echo "clean"
# 5. Verify telemetry landed in Grafana Cloud#    PromQL on the metrics datasource (should be > 0):#      sum(up{cluster="production-us-east"})#    LogQL on Loki:#      sum(count_over_time({cluster="production-us-east"}[5m]))

Full values.yaml, key PromQL, dashboard IDs (15520, 1860, 14282…), and alert rules: references/k8s-monitoring-values.md [blocked].

2. Monitor a Linux host

alloy
# 1. /etc/alloy/config.alloy — see references/clouds-and-hosts.md for the full blockprometheus.exporter.unix "host"  { rootfs_path = "/" }prometheus.scrape         "node" { targets = prometheus.exporter.unix.host.targets                                   forward_to = [prometheus.remote_write.cloud.receiver] }
bash
# 2. Reload Alloy and verify the unix exporter is upsystemctl reload alloycurl -s http://localhost:12345/api/v0/web/components | jq '.[] | select(.id|contains("prometheus.exporter.unix"))'
# 3. Verify in Grafana Cloud — open the "Node Exporter Full" dashboard (ID 1860)#    and pick your host from the `instance` dropdown.

3. Pull AWS / Azure / GCP metrics

Provision the datasource (full YAML in references/clouds-and-hosts.md [blocked]), then:

bash
# 1. After provisioning, restart Grafana to pick up the file# 2. Verify the datasource — Grafana → Connections → Data sources → "Test"#    Expect "Successfully queried the CloudWatch metrics API" (or equivalent).# 3. Confirm a query — Explore → datasource → metric e.g.#    CloudWatch namespace AWS/EC2 metric CPUUtilization, last 1h.

Troubleshooting

  • chart installed but no metrics in Cloud → check the grafana-cloud-secret api-key value; check Alloy logs for 401
  • kube-state-metrics pod Pending → likely RBAC; reapply the chart's CRDs/CRBs
  • Node-exporter pod CrashLoopBackOff → typically hostNetwork: true collision with the host's :9100; change the port
  • CloudWatch "Access denied" → IAM role missing cloudwatch:GetMetricData, cloudwatch:ListMetrics

Resources

來源與署名

來源:grafana/skills位於skills/grafana-cloud/infrastructure提交1ccacf2

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 grafana/skills 的技能

React 19 Plugin Migration

grafana

指導將 Grafana 外掛遷移至 React 19 相容,依序完成建置、相依性與原始碼修改步驟。

Software Development279今天更新

Plugin Bundle Size

grafana

指導使用 React.lazy、Suspense 與 webpack 程式碼分割來最佳化 Grafana 應用程式外掛的打包體積。

Software Development279今天更新

Grafana Scenes

grafana

使用 @grafana/scenes 框架建置 Grafana 外掛頁面,涵蓋場景、面板、變數與下鑽導覽。

Software Development279今天更新

Check Npm

grafana

對 JS/TS 儲存庫的 npm、yarn 或 pnpm 設定進行唯讀供應鏈強化稽核。

Security279今天更新

Mimir

grafana

指導架設與維運 Grafana Mimir,用於可擴充、多租戶、長期的 Prometheus 與 OTLP 指標儲存。

DevOps & Cloud279今天更新

K6 Trend Analysis

grafana

Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for performance regression", "should I tighten thresholds", "are my tests degrading", "show me trends for test X", "analyze my k6 test runs", or "is my test getting slower". Also trigger when a user asks to check all tests in a project -- run this skill once per test and synthesize.

待分類279今天更新