深色模式
资源利用率提升
大部分集群的真实装箱率在 20%~40%,也就是说付了 100 台的钱只用了 30 台。提升利用率是成本优化中收益最大的一项。
适用环境
- Kubernetes 集群 ≥ 3 个节点
- 已部署 metrics-server 与 Prometheus
- 服务可设置 resource requests/limits
bash
kubectl top nodes
kubectl describe nodes | grep -A5 'Allocated resources'1
2
2
操作步骤
1. 测量当前装箱率
bash
# 节点维度的分配率 vs 使用率
kubectl get nodes -o json \
| jq -r '.items[] | "\(.metadata.name) cpu_alloc=\(.status.allocatable.cpu) mem_alloc=\(.status.allocatable.memory)"'1
2
3
2
3
promql
# 真实 CPU 使用率(分配率≠使用率)
sum(rate(container_cpu_usage_seconds_total{container!=""}[5m]))
/ sum(kube_pod_container_resource_requests{resource="cpu"})1
2
3
2
3
text
装箱率 = 已分配 requests / 可分配总量
利用率 = 实际使用 / 已分配 requests
有效率 = 装箱率 × 利用率 = 实际使用 / 总量 ← 真正要提升的指标1
2
3
2
3
2. 找出"虚高 requests"
promql
# requests 与实际使用的比值,> 3 说明严重虚高
sum(kube_pod_container_resource_requests{resource="cpu"}) by (namespace, pod)
/ clamp_min(quantile_over_time(0.95,
sum(rate(container_cpu_usage_seconds_total[5m])) by (namespace, pod)[7d:5m]), 0.01)1
2
3
4
2
3
4
bash
# 输出 Top 20 优化候选
kubectl get deploy -A -o json \
| jq -r '.items[] | .metadata.namespace + "/" + .metadata.name + " " +
(.spec.template.spec.containers[] | "cpu=" + (.resources.requests.cpu // "-") + " mem=" + (.resources.requests.memory // "-"))' \
| sort | head -201
2
3
4
5
2
3
4
5
3. 修正 requests(分批进行)
yaml
resources:
requests:
cpu: "500m" # 按 P95 实测 × 1.2 设置
memory: "1Gi" # 按峰值 × 1.2 设置
limits:
memory: "2Gi" # 内存 limits 必须设
# cpu limits 视语言/延迟要求决定,Java 类常省略以避免 throttling1
2
3
4
5
6
7
2
3
4
5
6
7
bash
# 用 VPA 的推荐值作为参考(先只开 recommend 模式)
kubectl apply -f - <<'EOF'
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: order-api-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: order-api
updatePolicy:
updateMode: "Off" # 只给建议,不自动改
EOF
kubectl get vpa order-api-vpa -o jsonpath='{.status.recommendation.containerRecommendations}' | jq .1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
2
3
4
5
6
7
8
9
10
11
12
13
14
15
不要直接开启 VPA 自动模式
updateMode: Auto 会重启 Pod 来应用新规格,对在线服务会造成抖动。先用 Off 收集建议,人工确认后再分批应用。
4. 混部:把在线与批处理放一起
text
在线服务:白天高、夜间低,延迟敏感
批处理任务:可延迟执行,可抢占
夜间把批处理调度上去 → 利用在线服务空出的资源1
2
3
4
2
3
4
yaml
# 批处理设置低优先级,资源紧张时先被挤掉
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: batch-low
value: 100
preemptionPolicy: Never
---
apiVersion: batch/v1
kind: Job
metadata:
name: nightly-etl
spec:
template:
spec:
priorityClassName: batch-low
containers:
- name: etl
image: etl:latest
resources:
requests: {cpu: "4", memory: "8Gi"}
restartPolicy: OnFailure1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
5. 处理碎片与节点伸缩
碎片问题:每个节点都剩一点资源,但没有任何单节点能放下新 Pod,导致集群"看起来有资源却调度不进去"。用 descheduler 重调度整理:
bash
helm repo add descheduler https://kubernetes-sigs.github.io/descheduler/
helm install descheduler descheduler/descheduler -n kube-system \
--set deschedulerPolicy.profiles[0].plugins.balance.enabled[0]=RemoveDuplicates1
2
3
2
3
节点自动伸缩关键参数:
text
--scale-down-utilization-threshold=0.5 # 节点利用率低于 50% 可缩
--scale-down-unneeded-time=10m # 持续 10 分钟才可缩,防抖
节点选型:
大节点 → 装箱率高,但故障影响面大、碎片粒度粗
小节点 → 灵活,但系统组件开销占比高(可达 15%~25%)
推荐混合:核心服务中等规格,批处理大规格1
2
3
4
5
6
7
2
3
4
5
6
7
7. 计算收益
text
优化前:有效率 30%,100 节点
优化后:
requests 虚高修正 → 使用率从 25% 提到 60%
混部 + 碎片整理 → 装箱率从 60% 提到 80%
有效率 = 0.6 × 0.8 = 48%
同样业务量所需节点 ≈ 100 × 30 / 48 ≈ 63 节点1
2
3
4
5
6
7
2
3
4
5
6
7
验证
bash
# 1) 有效率是否提升
kubectl top nodes
# 2) 是否有 Pod 因资源不足 Pending
kubectl get pods -A --field-selector status.phase=Pending
# 3) throttling 与 OOM 是否增加
kubectl get events -A --field-selector reason=OOMKilling
# 4) P99 延迟是否恶化1
2
3
4
5
6
7
2
3
4
5
6
7
常见坑
只看分配率不看使用率
kubectl describe node 显示分配了 90%,但真实使用可能只有 20%。必须同时看使用率。
把内存 requests 调太低
内存是不可压缩资源,requests 低于实际需要会导致节点内存紧张时整机 OOM。内存 requests 必须基于峰值而非均值。
碎片被忽略
节点平均利用率 40% 但 Pod 调度不进去,是碎片问题而非容量不足。需要 descheduler 或统一 Pod 规格档位。
系统组件占用未计入
DaemonSet(日志、监控、网络插件)在每个节点都占资源。小节点集群中这部分占比可达 15%~25%,计算可用容量时必须扣除。