深色模式
Federation 联邦聚合
当机器和 Prometheus 太多,建一个“总控”Prometheus 从各分片拉取汇总指标,这就是联邦。
适用环境
- 已有多个按机房/业务分片的 Prometheus
- 需要统一视图与跨分片告警
bash
# 在分片 Prometheus 上确认可被联邦的接口
curl -s 'http://shard1:9090/federate?match[]=up' | head1
2
2
操作步骤
1. 下层暴露联邦接口
下层 Prometheus 无需特殊配置,/federate 接口默认可用,靠 match[] 选择要暴露的指标。
2. 上层配置 federation 抓取
yaml
- job_name: federate
scrape_interval: 30s
honor_labels: true
metrics_path: /federate
params:
match[]:
- '{job="node"}' # 只拉取需要的指标
- '{__name__=~"node_cpu.*"}'
static_configs:
- targets:
- 'shard1:9090'
- 'shard2:9090'1
2
3
4
5
6
7
8
9
10
11
12
2
3
4
5
6
7
8
9
10
11
12
3. 热加载
bash
curl -X POST http://localhost:9090/-/reload1
验证
promql
# 上层应能查到来自分片的聚合数据
count by (job) (up)1
2
2
常见坑
只联邦聚合后的指标
联邦不是“全量同步”,应只 match 关键/汇总指标,否则上层存储会被打爆。
honor_labels 必须开
不开 honor_labels: true 时,上层会用自己的 job/instance 覆盖原始标签,溯源丢失。