无需修改DeploymentConfig,如何将Pod调度至指定OpenShift GPU节点?
你的需求很明确:GPU节点只跑申请了nvidia.com/gpu资源的ML Pod,用户只要调整资源限制,完全不用碰亲和性或污点配置。这个方案完美匹配你的诉求,核心思路是用污点把GPU节点“锁”起来,再通过自动注入容忍度,让申请了GPU资源的Pod自动获得准入权限。
第一步:给GPU节点加“锁”(添加污点)
先给所有配备Tesla P40的节点添加污点,默认情况下所有普通Pod都无法调度到这些节点:
# 单个节点操作,替换<gpu-node-name>为你的节点名称 oc adm taint node <gpu-node-name> nvidia.com/gpu=required:NoSchedule # 批量操作(如果你的GPU节点有统一标识标签,比如nvidia.com/gpu.present=true) oc get nodes -l nvidia.com/gpu.present=true --no-headers | awk '{print $1}' | xargs oc adm taint node nvidia.com/gpu=required:NoSchedule
第二步:自动给GPU Pod开“钥匙”(部署Mutating Webhook)
我们需要一个Mutating Admission Webhook,它会在Pod创建时自动检查:如果Pod的资源限制里包含nvidia.com/gpu,就自动给Pod添加容忍刚才那个污点的配置。用户完全不用手动加容忍度,只要添加GPU资源限制就行。
2.1 部署Webhook服务
你可以用Go/Python写一个简单的Webhook服务,核心逻辑就是:检查Pod的spec.containers[*].resources.limits里有没有nvidia.com/gpu,有的话就往spec.tolerations里插入一行:
- key: "nvidia.com/gpu" operator: "Equal" value: "required" effect: "NoSchedule"
把这个服务打包成镜像,部署到OpenShift的运维命名空间(比如openshift-cluster-node-tuning-operator),并创建对应Service:
apiVersion: apps/v1 kind: Deployment metadata: name: gpu-toleration-webhook namespace: openshift-cluster-node-tuning-operator spec: replicas: 1 selector: matchLabels: app: gpu-toleration-webhook template: metadata: labels: app: gpu-toleration-webhook spec: containers: - name: webhook image: <你的Webhook镜像地址> ports: - containerPort: 443 --- apiVersion: v1 kind: Service metadata: name: gpu-toleration-webhook-svc namespace: openshift-cluster-node-tuning-operator spec: selector: app: gpu-toleration-webhook ports: - port: 443 targetPort: 443
2.2 注册Webhook到Kubernetes
创建MutatingWebhookConfiguration,让Kubernetes在创建Pod时自动调用这个Webhook:
apiVersion: admissionregistration.k8s.io/v1 kind: MutatingWebhookConfiguration metadata: name: gpu-toleration-webhook webhooks: - name: gpu-toleration.yourdomain.com clientConfig: service: name: gpu-toleration-webhook-svc namespace: openshift-cluster-node-tuning-operator path: /mutate caBundle: <你的CA证书Base64编码内容> # 需生成并配置证书,确保Webhook被Kubernetes信任 rules: - apiGroups: [""] apiVersions: ["v1"] operations: ["CREATE"] resources: ["pods"] sideEffects: None admissionReviewVersions: ["v1"] # 可选:排除系统命名空间,避免影响系统组件Pod namespaceSelector: matchExpressions: - key: name operator: NotIn values: ["kube-system", "openshift-system"]
第三步:验证效果
测试一下就能确认方案是否生效:
- 创建一个普通的DeploymentConfig(无GPU资源限制):
apiVersion: apps.openshift.io/v1 kind: DeploymentConfig metadata: name: test-normal-app spec: replicas: 1 selector: app: test-normal template: metadata: labels: app: test-normal spec: containers: - name: nginx image: nginx resources: limits: cpu: "1"
这个Pod会被污点阻挡,无法调度到GPU节点。
- 创建一个带GPU资源限制的ML DeploymentConfig:
apiVersion: apps.openshift.io/v1 kind: DeploymentConfig metadata: name: test-ml-gpu-app spec: replicas: 1 selector: app: test-ml-gpu template: metadata: labels: app: test-ml-gpu spec: containers: - name: ml-worker image: nvidia/cuda:11.8.0-runtime-ubuntu22.04 resources: limits: nvidia.com/gpu: 1
这个Pod会被Webhook自动添加容忍度,顺利调度到GPU节点,用户全程只修改了资源限制,完全没碰亲和性或污点配置。
备选方案:自定义调度器Predicate(不推荐,复杂度高)
如果你不想用Webhook,也可以自定义OpenShift调度器的Predicate函数,逻辑是:只有Pod请求了nvidia.com/gpu资源,才允许调度到带GPU标签的节点。但这个方法需要修改调度器代码、重新编译部署,复杂度很高,除非有特殊需求,否则不建议使用。
内容的提问来源于stack exchange,提问作者白栋天

