Infrastructure

作者 grafana1ccacf29049fApache-2.0279 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Ship Kubernetes, host, container, and cloud-provider telemetry into Grafana Cloud — `k8s-monitoring` Helm chart for K8s clusters (metrics + logs + traces + events + cost), Alloy `prometheus.exporter.unix` for Linux hosts, cAdvisor + Docker discovery for containers, and CloudWatch / Azure Monitor / Google Cloud Monitoring datasource setup. Use when onboarding a new cluster or VM fleet to Grafana Cloud, picking the right Helm values for K8s scraping, wiring kube-state-metrics + node-exporter + cAdvisor, alerting on `PodCrashLooping` / node memory / PVC capacity, or pulling AWS / Azure / GCP cloud metrics — even when the user says "monitor my cluster", "send K8s metrics to Grafana", "scrape EC2 metrics", "cluster pod logs", or "install the monitoring helm chart" without naming `k8s-monitoring` or Alloy.

仅含说明DevOps & Cloud
AI 生成的概览

使用 Helm 和 Alloy 将 Kubernetes 集群、Linux 主机、容器和云指标接入 Grafana Cloud。

功能
该技能提供将 Kubernetes、主机、容器和云提供商遥测数据发送到 Grafana Cloud 的分步说明。涵盖安装 grafana/k8s-monitoring Helm chart、为 Linux 主机配置 Alloy 的 prometheus.exporter.unix,以及配置 CloudWatch、Azure Monitor 或 Google Cloud Monitoring 数据源。还包含验证命令和常见问题的故障排除指南。
适用场景
适用于将新的 Kubernetes 集群或虚拟机集群接入 Grafana Cloud、为 K8s 抓取选择 Helm values,或配置 kube-state-metrics、node-exporter 和 cAdvisor。也适合设置 PodCrashLooping、节点内存或 PVC 容量告警,以及拉取 AWS、Azure 或 GCP 云指标。
运行要求
需要具有 Prometheus、Loki 和 Tempo 端点以及具备 metrics:write、logs:write 和 traces:write 权限的 API 密钥的 Grafana Cloud 堆栈。对于 Kubernetes,需要集群以及 helm 3.x 和 kubectl 上下文。对于主机或 Docker,需要在节点上安装 Alloy。不附带脚本;仅为说明。

Grafana Cloud Infrastructure Monitoring

Docs: https://grafana.com/docs/grafana-cloud/monitor-infrastructure/

K8s + host + container + cloud-provider telemetry, mostly via the grafana/k8s-monitoring Helm chart or Alloy.

Prerequisites

  • Grafana Cloud stack with Prometheus / Loki / Tempo endpoints + API key (metrics:write, logs:write, traces:write)
  • For Kubernetes: a cluster + helm 3.x + kubectl context pointing at it
  • For hosts / Docker: Alloy installed on the node

Common Workflows

1. Onboard a Kubernetes cluster (k8s-monitoring chart)

bash
# 1. Create the namespace + secretkubectl create namespace monitoringkubectl create secret generic grafana-cloud-secret \  -n monitoring --from-literal=api-key=<your-api-key>
# 2. Install — values.yaml in references/k8s-monitoring-values.mdhelm repo add grafana https://grafana.github.io/helm-charts && helm repo updatehelm install k8s-monitoring grafana/k8s-monitoring \  --version 4.1.4 -n monitoring -f values.yaml
# 3. Verify every pod is Runningkubectl get pods -n monitoring# Expect alloy-*, kube-state-metrics-*, node-exporter-*, etc. all Ready.
# 4. Verify no error logs in the metrics/logs/traces Alloyskubectl -n monitoring logs deploy/k8s-monitoring-alloy-metrics --tail=50 | grep -iE 'error|level=err' || echo "clean"
# 5. Verify telemetry landed in Grafana Cloud#    PromQL on the metrics datasource (should be > 0):#      sum(up{cluster="production-us-east"})#    LogQL on Loki:#      sum(count_over_time({cluster="production-us-east"}[5m]))

Full values.yaml, key PromQL, dashboard IDs (15520, 1860, 14282…), and alert rules: references/k8s-monitoring-values.md [blocked].

2. Monitor a Linux host

alloy
# 1. /etc/alloy/config.alloy — see references/clouds-and-hosts.md for the full blockprometheus.exporter.unix "host"  { rootfs_path = "/" }prometheus.scrape         "node" { targets = prometheus.exporter.unix.host.targets                                   forward_to = [prometheus.remote_write.cloud.receiver] }
bash
# 2. Reload Alloy and verify the unix exporter is upsystemctl reload alloycurl -s http://localhost:12345/api/v0/web/components | jq '.[] | select(.id|contains("prometheus.exporter.unix"))'
# 3. Verify in Grafana Cloud — open the "Node Exporter Full" dashboard (ID 1860)#    and pick your host from the `instance` dropdown.

3. Pull AWS / Azure / GCP metrics

Provision the datasource (full YAML in references/clouds-and-hosts.md [blocked]), then:

bash
# 1. After provisioning, restart Grafana to pick up the file# 2. Verify the datasource — Grafana → Connections → Data sources → "Test"#    Expect "Successfully queried the CloudWatch metrics API" (or equivalent).# 3. Confirm a query — Explore → datasource → metric e.g.#    CloudWatch namespace AWS/EC2 metric CPUUtilization, last 1h.

Troubleshooting

  • chart installed but no metrics in Cloud → check the grafana-cloud-secret api-key value; check Alloy logs for 401
  • kube-state-metrics pod Pending → likely RBAC; reapply the chart's CRDs/CRBs
  • Node-exporter pod CrashLoopBackOff → typically hostNetwork: true collision with the host's :9100; change the port
  • CloudWatch "Access denied" → IAM role missing cloudwatch:GetMetricData, cloudwatch:ListMetrics

Resources

来源与署名

来源:grafana/skills位于skills/grafana-cloud/infrastructure提交1ccacf2

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 grafana/skills 的技能

React 19 Plugin Migration

grafana

指导将 Grafana 插件迁移至 React 19 兼容,按顺序完成构建、依赖与源码修改步骤。

Software Development279今天更新

Plugin Bundle Size

grafana

指导使用 React.lazy、Suspense 和 webpack 代码分割来优化 Grafana 应用插件包体积。

Software Development279今天更新

Grafana Scenes

grafana

使用 @grafana/scenes 框架构建 Grafana 插件页面,涵盖场景、面板、变量与下钻导航。

Software Development279今天更新

Check Npm

grafana

对 JS/TS 仓库的 npm、yarn 或 pnpm 配置进行只读供应链加固审计。

Security279今天更新

Mimir

grafana

指导搭建和运维 Grafana Mimir,用于可扩展、多租户、长期的 Prometheus 与 OTLP 指标存储。

DevOps & Cloud279今天更新

K6 Trend Analysis

grafana

Analyze Grafana Cloud k6 test run trends over time. Detects slow metric drift (e.g., P95 latency creeping up while still passing thresholds), computes headroom to thresholds, flags anomalies, and recommends threshold tightening. Use when the user asks about test performance trends, wants to know if metrics are degrading, asks whether thresholds should be tightened, or wants a health check across recent runs for a specific test. Trigger on phrases like "how is my test trending", "is P95 getting worse", "check for performance regression", "should I tighten thresholds", "are my tests degrading", "show me trends for test X", "analyze my k6 test runs", or "is my test getting slower". Also trigger when a user asks to check all tests in a project -- run this skill once per test and synthesize.

待分类279今天更新