深色模式
HPA 自动伸缩实践
摘要:本文从零配置一个可用的 HPA:先压测看效果,再解释平均利用率的算法、扩容与缩容的冷却时间、多指标取最大值的策略,以及 HPA 显示 unknown 的排查办法。
适用环境
- 可用 K8s 集群
- 必须已安装 metrics-server(否则 HPA 拿不到指标)
- 有一个 Deployment,且设置了 resources.requests
操作步骤
一、安装 metrics-server
bash
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
kubectl -n kube-system get pod -l k8s-app=metrics-server
kubectl top nodes
kubectl top pods1
2
3
4
2
3
4
若 Pod 起不来(常见于自签证书环境):
bash
kubectl -n kube-system patch deploy metrics-server --type=json \
-p='[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'1
2
2
注意
--kubelet-insecure-tls 跳过证书校验,仅用于测试环境。生产应配置正确的 CA 与 kubelet 证书。
二、准备带资源请求的 Deployment
yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: php-apache
spec:
replicas: 1
selector:
matchLabels:
app: php-apache
template:
metadata:
labels:
app: php-apache
spec:
containers:
- name: php-apache
image: registry.k8s.io/hpa-example
ports:
- containerPort: 80
resources:
requests:
cpu: 200m
memory: 128Mi
limits:
cpu: 500m
memory: 256Mi
---
apiVersion: v1
kind: Service
metadata:
name: php-apache
spec:
selector:
app: php-apache
ports:
- port: 801
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
危险
没有 resources.requests 的 Pod,HPA 无法计算利用率,会一直显示 unknown 且永不扩容。 这是 HPA 不生效的头号原因。
三、创建 HPA
bash
kubectl autoscale deployment php-apache \
--cpu-percent=50 --min=1 --max=101
2
2
或用 YAML(可版本化,推荐):
yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: php-apache
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: php-apache
minReplicas: 1
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 70
behavior:
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 15
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 50
periodSeconds: 601
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
四、压测看效果
bash
kubectl run load --rm -it --image=busybox --restart=Never -- \
sh -c "while true; do wget -q -O- http://php-apache; done"1
2
2
另开一个终端观察:
bash
kubectl get hpa php-apache -w
kubectl get pod -l app=php-apache -w1
2
2
五、理解扩容算法
期望副本数 = ceil(当前副本数 × 当前平均利用率 / 目标利用率)。例如目标 50%、当前 100%、副本 1 → 期望 2。多个指标时取计算结果的最大值,即最激进的那个指标说了算。
六、冷却机制
- 扩容:默认无稳定窗口,快速扩,但 3 分钟内不会再次扩容。
- 缩容:默认
stabilizationWindowSeconds: 300,即指标下降后要持续观察 5 分钟才缩,避免抖动。 - 通过
behavior字段可自定义(见上面 YAML)。
七、自定义指标(QPS)
基于 QPS 需要额外组件(Prometheus Adapter)提供 pods 类型自定义指标:
yaml
metrics:
- type: Pods
pods:
metric:
name: http_requests_per_second
target:
type: AverageValue
averageValue: "100"1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
建议
CPU 指标简单但滞后;QPS 更贴近业务但依赖监控链路。生产建议 CPU 打底 + 自定义业务指标辅助。
验证
- [ ]
kubectl top pods能返回实际数值(不是<unknown>) - [ ] 压测后
kubectl get hpa的 TARGETS 从0%/50%上升 - [ ] 副本数随负载增加到 max 上限
- [ ] 停止压测 5 分钟后副本数回落
常见坑
- TARGETS 显示
<unknown>/50%:metrics-server 未装/异常,或 Pod 没设resources.requests。 - HPA 有数值但副本不变:当前利用率已低于目标,或未到冷却时间。
- 副本疯狂上下抖动:缩容稳定窗口太短,把
scaleDown.stabilizationWindowSeconds调大。 - 扩容了但服务仍然慢:瓶颈在数据库等下游,扩应用无用。
- max 太小:
maxReplicas设成 2,压测时撑不住,需按峰值流量预留。