深色模式
K8s 监控指标:控制面、节点与工作负载
摘要:本文按控制面、节点、工作负载三层梳理该看哪些指标,演示 metrics-server 与 kube-state-metrics 的部署,给出 Prometheus 抓取配置与几条真正有用的告警规则。
适用环境
- 可用 K8s 集群 +
kubectl - 建议先装 metrics-server(基础),进阶再上 Prometheus
- 有 Helm 便于部署 Prometheus
操作步骤
一、先明确三层监控对象
| 层级 | 关注什么 | 典型指标 |
|---|---|---|
| 控制面 | apiserver 可用性与延迟、etcd 健康、调度是否卡 | apiserver 请求延迟、etcd 写入延迟 |
| 节点 | CPU/内存/磁盘/网络资源饱和度 | node_cpu、node_memory、disk IO |
| 工作负载 | Pod 是否就绪、是否重启、资源用量 | 容器 CPU/内存、restart 次数 |
二、基础层:metrics-server
bash
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
kubectl top nodes
kubectl top pods -A --sort-by=memory | head -20
kubectl top pod --containers -n <命名空间>1
2
3
4
2
3
4
它只提供当前瞬时值,不保存历史,供 kubectl top 和 HPA 使用。
三、kube-state-metrics:对象状态指标
metrics-server 看不到「Deployment 期望 3 但实际 1」这类状态,kube-state-metrics 补齐这一块:
bash
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm install kube-state-metrics prometheus-community/kube-state-metrics -n monitoring --create-namespace
kubectl -n monitoring get pod1
2
3
2
3
它生成的指标示例:
text
kube_deployment_status_replicas_ready
kube_deployment_spec_replicas
kube_pod_status_phase
kube_pod_container_status_restarts_total
kube_node_status_condition1
2
3
4
5
2
3
4
5
四、部署 Prometheus(kube-prometheus-stack)
bash
helm install prometheus prometheus-community/kube-prometheus-stack \
-n monitoring --create-namespace \
--set prometheus.prometheusSpec.retention=15d
kubectl -n monitoring get pod,svc
kubectl -n monitoring port-forward svc/prometheus-kube-prometheus-prometheus 9090:90901
2
3
4
5
2
3
4
5
访问 http://localhost:9090 进入 PromQL 界面。
五、控制面关键指标
promql
# apiserver 请求延迟 P99
histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket[5m])) by (le, verb))
# apiserver 错误率
sum(rate(apiserver_request_total{code=~"5.."}[5m])) / sum(rate(apiserver_request_total[5m]))
# etcd 写入延迟(超过 100ms 要警惕)
histogram_quantile(0.99, sum(rate(etcd_request_duration_seconds_bucket{operation="put"}[5m])) by (le))
# 调度失败
sum(rate(scheduler_schedule_attempts_total{result="unschedulable"}[5m]))
# etcd 剩余容量(默认配额 2GB)
etcd_mvcc_db_total_size_in_bytes / etcd_server_quota_backend_bytes1
2
3
4
5
6
7
8
9
10
11
12
13
14
2
3
4
5
6
7
8
9
10
11
12
13
14
六、节点关键指标(USE 方法)
promql
# CPU 使用率
100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance) * 100)
# 内存可用率
node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100
# 磁盘将在 4 小时内写满(基于 6 小时增长率)
predict_linear(node_filesystem_free_bytes[6h], 4*3600) < 0
# 磁盘 IO 等待
rate(node_disk_io_time_seconds_total[5m]) * 1001
2
3
4
5
6
7
8
9
10
11
2
3
4
5
6
7
8
9
10
11
七、工作负载关键指标(RED 方法)
promql
# 容器重启次数(10 分钟内)
increase(kube_pod_container_status_restarts_total[10m]) > 0
# Pod 未就绪
kube_deployment_status_replicas_ready / kube_deployment_spec_replicas < 1
# 容器 CPU 使用率(相对 requests)
sum(rate(container_cpu_usage_seconds_total{pod!=""}[5m])) by (pod)
/ sum(kube_pod_container_resource_requests{resource="cpu"}) by (pod)
# 内存接近 limits(超过 85% 告警)
container_memory_working_set_bytes{pod!=""}
/ kube_pod_container_resource_limits{resource="memory"} > 0.85
# CPU 被限流的容器(throttling)
rate(container_cpu_cfs_throttled_seconds_total[5m]) > 0.11
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
建议
上面这条 CPU throttling 指标非常有价值。服务莫名变慢却看不出原因时,多半是被限流,而 kubectl top 完全看不出来。
八、几条真正有用的告警规则
yaml
groups:
- name: k8s.rules
rules:
- alert: NodeNotReady
expr: kube_node_status_condition{condition="Ready",status="true"} == 0
for: 5m
annotations:
summary: "节点 {{ $labels.node }} 已 NotReady 超过 5 分钟"
- alert: PodCrashLooping
expr: increase(kube_pod_container_status_restarts_total[15m]) > 3
for: 2m
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.pod }} 频繁重启"
- alert: DeploymentReplicasMismatch
expr: kube_deployment_status_replicas_ready != kube_deployment_spec_replicas
for: 10m
annotations:
summary: "Deployment {{ $labels.deployment }} 副本数未达期望"
- alert: EtcdDiskWillFill
expr: etcd_mvcc_db_total_size_in_bytes / etcd_server_quota_backend_bytes > 0.85
for: 5m
annotations:
summary: "etcd 数据库已用超过 85% 配额"1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
九、验证指标链路
bash
kubectl -n monitoring get pod
kubectl -n monitoring port-forward svc/prometheus-operated 9090:9090
# 浏览器打开后执行:up1
2
3
2
3
up 指标为 1 说明抓取正常,为 0 说明 target 有问题,去 Prometheus 的 Status → Targets 看具体错误。
验证
- [ ]
kubectl top nodes返回真实数值 - [ ] Prometheus 的 Targets 页面全部 UP
- [ ] 能查出容器 CPU/内存用量曲线
- [ ] 人为重启一个 Pod,PodCrashLooping 类告警能触发
常见坑
kubectl top报metrics not available:metrics-server 未装或启动失败(常见于证书问题,需--kubelet-insecure-tls)。- Prometheus 抓不到控制面指标:apiserver、etcd、scheduler 的 metrics 端口在 host 网络或被 NetworkPolicy 拦截。
- 指标有数据但告警不触发:
for时长设置与查询窗口不匹配,先在 Prometheus 界面手动验证表达式有结果。 - 只监控资源不监控状态:资源正常但 Deployment 副本为 0 的情况完全看不出来,必须装 kube-state-metrics。
- 存储规划不足:Prometheus 数据增长快,retention 与磁盘容量要提前算好,否则磁盘写满后采集停止。