深色模式
Ray 分布式计算与训练
摘要:本文面向需要在 Kubernetes 上运行分布式 Python / 训练 / 批处理 / Ray Serve 的 SRE / 平台工程师。讲解 KubeRay 的三种 CRD(
RayCluster/RayJob/RayService)、head 与 worker 的拓扑、GPU 训练实践、内置自动扩缩,以及常见故障排查。适用版本:KubeRay v1.6/v1.7([版本相关:以官方 release 为准]),Ray 运行时镜像 2.31.0([版本相关]),apiVersionray.io/v1。
适用版本与前提
- Kubernetes:v1.28+
- 已安装 KubeRay operator(
kubectl apply -k官方 manifests 或 Helmkuberay/ray-operator) - GPU 节点池就绪(训练场景),见 gpu-scheduling.md
- 对象存储(S3/MinIO)用于落地训练产物与 History Server 日志
核心概念:KubeRay 的三种资源
| CRD | 生命周期 | 典型用途 |
|---|---|---|
RayCluster | 长驻集群(1 head + N worker 组) | 交互式开发、常驻 Ray Serve |
RayJob | 运行到完成,可自动销毁集群 | 分布式训练、RL 后训练、批处理 |
RayService | RayCluster + Serve 配置,零停机升级 | 在线推理服务(HA) |
一个常见误区:不要在生产裸跑 ray start --head,那会绕开 K8s 调度、配额与配额隔离。KubeRay 让 Ray Pod 成为一等公民,GPU Operator、Kueue 配额、节点 taint 都能像对待普通负载一样约束它。
生产实践 1:声明式 RayCluster(含 GPU worker 组)
yaml
# ray-cluster.yaml
apiVersion: ray.io/v1
kind: RayCluster
metadata:
name: gpu-cluster
namespace: ray
spec:
rayVersion: "2.31.0" # [版本相关: 与镜像一致]
headGroupSpec:
rayStartParams:
dashboard-host: "0.0.0.0"
template:
spec:
containers:
- name: ray-head
image: rayproject/ray:2.31.0-py310 # [版本相关]
ports:
- containerPort: 6379 # GCS
- containerPort: 8265 # Dashboard
resources:
limits:
cpu: "4"
memory: 8Gi
requests:
cpu: "2"
memory: 4Gi
workerGroupSpecs:
- groupName: gpu-workers
replicas: 2
minReplicas: 1
maxReplicas: 8
rayStartParams:
num-gpus: "1"
template:
spec:
containers:
- name: ray-worker
image: rayproject/ray:2.31.0-py310
resources:
limits:
cpu: "8"
memory: 32Gi
nvidia.com/gpu: 1
requests:
cpu: "4"
memory: 16Gi
nvidia.com/gpu: 11
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
自动扩缩的前提
内置自动扩缩需在 RayCluster 开启 enableInTreeAutoscaling: true,且每个 workerGroupSpec 都必须设 minReplicas/maxReplicas,否则该标志虽为 true 也会被静默忽略。扩缩依据是 Ray 内 pending 的 task/actor(而非 CPU 利用率),对"请求 8 GPU 的 placement group"会自动拉起 8 个 GPU Pod。[未实测:行为以 KubeRay 文档为准]
生产实践 2:RayJob 跑分布式训练(运行即销毁)
yaml
# ray-job.yaml
apiVersion: ray.io/v1
kind: RayJob
metadata:
name: train-job
namespace: ray
spec:
entrypoint: "python train.py --epochs 5"
runtimeEnv: |
pip:
- "torch==2.3.0"
- "ray[train]==2.31.0"
shutdownAfterJobFinishes: true # 训练完自动回收集群
ttlSecondsAfterFinished: 300
rayClusterSpec:
rayVersion: "2.31.0"
headGroupSpec:
# 同 RayCluster head 配置(省略,参考上文)
rayStartParams:
dashboard-host: "0.0.0.0"
template:
spec:
containers:
- name: ray-head
image: rayproject/ray:2.31.0-py310
resources:
requests: { cpu: "2", memory: 4Gi }
limits: { cpu: "4", memory: 8Gi }
workerGroupSpecs:
- groupName: gpu-workers
replicas: 2
minReplicas: 1
maxReplicas: 8
rayStartParams:
num-gpus: "1"
template:
spec:
containers:
- name: ray-worker
image: rayproject/ray:2.31.0-py310
resources:
limits:
nvidia.com/gpu: 1
memory: 32Gi
requests:
nvidia.com/gpu: 1
memory: 16Gi1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
训练脚本 train.py 用 Ray Train 做数据并行:
python
# train.py —— 示意,需在自有环境实测
import ray
from ray import train
from ray.train.torch import TorchTrainer
def train_func():
# 此处放置真实训练逻辑(模型/优化器/数据加载)
for step in range(5):
train.report({"loss": 1.0 / (step + 1)}) # 上报指标
ray.init()
trainer = TorchTrainer(train_func, scaling_config=train.ScalingConfig(num_workers=2, use_gpu=True))
result = trainer.fit()
print(result.metrics)1
2
3
4
5
6
7
8
9
10
11
12
13
14
2
3
4
5
6
7
8
9
10
11
12
13
14
镜像与版本
Ray 镜像 tag 与 rayVersion、runtimeEnv 中的 ray[train] 版本需一致,否则 GCS 握手或序列化会失败。脚本仅展示接口形态,未含真实模型逻辑,[未实测]。
验证
bash
kubectl apply -f ray-cluster.yaml
kubectl get raycluster -n ray
# 期望: NAME STATUS AGE
# gpu-cluster Ready 2m
# 进入 head 提交任务
kubectl exec -it gpu-cluster-head-0 -n ray -- ray status
# 期望显示 head + worker 数量与资源(示意):
# ===== Active nodes =====
# Nodes: 3 (head + 2 workers)
# 查看 Ray Dashboard(端口转发)
kubectl -n ray port-forward svc/gpu-cluster-head-svc 8265:82651
2
3
4
5
6
7
8
9
10
11
12
13
2
3
4
5
6
7
8
9
10
11
12
13
回滚与清理
bash
# RayJob 完成后(shutdownAfterJobFinishes=true)集群自动销毁;如需提前清理:
kubectl delete rayjob train-job -n ray
# 长驻集群按需删除:
kubectl delete raycluster gpu-cluster -n ray1
2
3
4
2
3
4
删除即释放算力
删除 RayCluster/J 会立即释放 GPU,正在运行的训练会中断且不自动保存中间产物。训练务必把 checkpoint 落到对象存储,shutdownAfterJobFinishes 前确认产物已上传。
故障排查
| 现象 | 原因 | 排查 |
|---|---|---|
Worker 一直 Pending | 无满足 nvidia.com/gpu 的节点 / taint 未容忍 | kubectl describe pod 看 FailedScheduling |
| Head 起不来 | rayVersion 与镜像不一致 | 对一遍镜像 tag 与 spec.rayVersion |
| 自动扩缩不触发 | 未开 enableInTreeAutoscaling 或 worker 组缺 min/max | 查 operator 日志与 CRD 校验 |
| 对象存储放不下 | object-store-memory 过小 | 调大 rayStartParams.object-store-memory(字节) |
安全与合规
多租户与数据泄露
- 隔离:不同团队用独立
RayCluster+ 独立 namespace;用 ResourceQuota 与 LimitRange 限制 GPU 上限,防单团队吃满整池。 - 网络:KubeRay v1.7 起支持原生 NetworkPolicy 与 mTLS([版本相关,需对应 CRD/特性门控]);跨 Pod 通信默认在同一集群内,避免把 Ray Dashboard 公网暴露。
- 凭证:训练脚本读取对象存储的 AK/SK 走 Secret 注入,禁止硬编码;收敛 egress,防止训练数据外传。
- Dashboard 鉴权:启用 Ray token 认证(
spec.authOptions,v1.5.1+ 引入,[版本相关]),否则 Dashboard 可读任务与日志。
成本与性能
- GPU 拓扑:多卡训练需 GPU 间 NVLink 高速互联,错误放置(跨 PCIe 而非 NVLink)可使通信密集训练慢 3–8 倍(第三方经验值,[未实测])。用 GPU Feature Discovery 打拓扑标签 + 亲和性约束。
- 成本示例([价格随云厂商/时点变化,未实测]):H100 × 8 节点约 2–3 USD/GPU·小时;一次 5 epoch 训练若耗时 3 小时、占 16 卡,约 16×2.5×3≈120 USD。用
maxReplicas限制峰值、Spot/抢占实例跑非关键训练可显著降本。 - gang 调度:分布式训练必须同时拿到全部 GPU,否则死锁——配合 Kueue/Volcano 的 PodGroup gang 调度(见 gpu-scheduling.md)。