深色模式
升级 Escalation 策略
摘要:升级策略的本质是"人没响应时系统自动替他找人",而不是等别人良心发现。本文给出分级升级时限表、Alertmanager 的 routes/receivers 配置写法,以及不打扰业务就能验证升级链的方法。
适用环境
yaml
# Prometheus Alertmanager:以下配置片段适用于 v0.25+
route:
receiver: 'primary'
group_by: ['alertname', 'service']
group_wait: 30s
group_interval: 2m
repeat_interval: 4h1
2
3
4
5
6
7
2
3
4
5
6
7
操作步骤
第 1 步:定升级时限
| 级别 | 首响要求 | 超时升级对象 | 再超时 |
|---|---|---|---|
| Sev1 | 5 分钟 | Secondary | 10 分钟 → 团队负责人 |
| Sev2 | 15 分钟 | Secondary | 30 分钟 → 团队负责人 |
| Sev3 | 30 分钟 | 次日处理 | 无需升级 |
第 2 步:用 route 树按级别分派
yaml
route:
receiver: 'default-catch-all'
routes:
- matchers:
- severity="critical"
receiver: 'oncall-primary'
group_wait: 0s
continue: false
- matchers:
- severity="warning"
receiver: 'oncall-secondary'
group_wait: 2m
receivers:
- name: 'oncall-primary'
webhook_configs:
- url: 'http://escalation-gateway/webhook'
send_resolved: true1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
第 3 步:在接收端做"未 ACK 即升级"
升级通常由值班平台(PagerDuty / Grafana OnCall / 自建网关)实现。自建时用一条规则:告警进入后若 N 分钟无 ACK 记录,则转发给下一级。
bash
# 自建网关的最小实现思路(伪生产代码,示意)
cat > escalation-gateway.sh <<'EOF'
#!/usr/bin/env bash
# 入参: $1=级别 $2=告警ID
set -euo pipefail
sev=$1; id=$2
ack_file=/var/run/ack/$id
deadline=$([ "$sev" = "critical" ] && echo 300 || echo 900)
for ((i=0; i<deadline; i+=10)); do
[ -f "$ack_file" ] && { echo "已 ACK,停止升级"; exit 0; }
sleep 10
done
echo "超时未 ACK,升级到下一级"
curl -sS -X POST "$NEXT_LEVEL_WEBHOOK" -d "{\"id\":\"$id\",\"sev\":\"$sev\"}"
EOF
chmod +x escalation-gateway.sh1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
第 4 步:演练升级链
bash
# 用一条名为 EscalationDrill 的告警发到非生产 receiver
amtool alert add EscalationDrill severity=critical service=drill \
--alertmanager.url=http://127.0.0.1:9093
# 查看是否被路由到正确的 receiver
amtool alert query severity=critical --alertmanager.url=http://127.0.0.1:90931
2
3
4
5
6
2
3
4
5
6
第 5 步:记录每次升级
bash
mkdir -p escalation-log
echo "$(date -u '+%F %T') sev=$sev id=$id -> next_level" >> escalation-log/$(date +%F).log1
2
2
验证
bash
# 1. 配置语法正确(必须在上线前跑)
amtool check-config /etc/alertmanager/alertmanager.yml
# 2. 路由树符合预期
amtool config routes test --config.file=/etc/alertmanager/alertmanager.yml \
severity=critical service=pay
# 3. 确认演练告警未误发到生产 channel
amtool alert query alertname=EscalationDrill --alertmanager.url=http://127.0.0.1:90931
2
3
4
5
6
7
8
9
2
3
4
5
6
7
8
9
常见坑
只有一级没有兜底
Primary 失联时若没有再上一级,告警会静默丢弃。升级链至少三层:Primary → Secondary → 负责人。
group_wait 拖慢首响
group_wait 默认 30s 会延迟聚合,critical 级别应设为 0s 立即发出。
升级即 P0 全员轰炸
把所有超时都升级为"叫醒所有人",会迅速耗尽团队信任。升级强度必须与事件级别匹配。