You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes节点亲和性与污点配置问题:Pod调度失败求助

Pod调度失败排查:Node Affinity与Taints配置问题

问题背景

测试Node Affinity与Taints功能,目标是将Pod调度到指定节点:

  • 目标节点已添加Label:node=testvm
  • 目标节点已添加Taint:node=testvm:NoSchedule

使用的Pod清单如下:

apiVersion: v1
kind: Pod
metadata:
  name: nginx
spec:
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:
        - matchExpressions:
          - key: node
            operator: In
            values:
            - testvm 
  containers:
  - name: nginx
    image: nginx
    imagePullPolicy: IfNotPresent

遇到的问题

1. 应用清单时的修改错误

执行应用操作后返回以下错误:

* spec.tolerations: Forbidden: existing toleration can not be modified except its tolerationSeconds
* spec: Forbidden: pod updates may not change fields other than `spec.containers[*].image`, `spec.initContainers[*].image`, `spec.activeDeadlineSeconds`, `spec.tolerations` (only additions to existing tolerations) or `spec.terminationGracePeriodSeconds` (allow it to be set to 1 if it was previously negative)
  core.PodSpec{
    ... // 15 identical fields
    Subdomain:         "",
    SetHostnameAsFQDN: nil,
-   Affinity: &core.Affinity{
-       NodeAffinity: &core.NodeAffinity{
-           RequiredDuringSchedulingIgnoredDuringExecution: &core.NodeSelector{NodeSelectorTerms: []core.NodeSelectorTerm{...}},
-       },
-   },

2. 调度警告

后续出现调度失败警告:

Warning  FailedScheduling   8s (x1 over 68s)  default-scheduler   0/4 nodes are available: 1 node(s) didn't match Pod's node affinity/selector, 1 node(s) had taint {nodepool: monitoring}, that the pod didn't tolerate, 1 node(s) had taint {nodepool: stage}, that the pod didn't tolerate, 1 node(s) had taint {nodepool: system}, that the pod didn't tolerate.
  Normal   NotTriggerScaleUp  2m23s             cluster-autoscaler  pod didn't trigger scale-up: 1 max node group size reached

原因分析

  1. Pod修改权限限制:Kubernetes不允许修改已存在Pod的affinity字段,这类调度核心配置只能在Pod创建时定义,无法通过更新操作修改。错误信息明确列出了允许修改的字段范围,affinity不在其中。
  2. 缺少Toleration配置:目标节点带有node=testvm:NoSchedule污点,当前Pod未配置对应的容忍规则,调度器会拒绝将Pod调度到该节点。同时其他节点要么带有Pod无法容忍的其他污点,要么不匹配Pod的亲和规则,导致无可用节点。

解决方法

1. 处理Pod修改错误

删除已存在的nginxPod,重新创建包含完整配置的Pod:

kubectl delete pod nginx

2. 修正Pod配置(添加Toleration)

更新Pod清单,添加对目标节点污点的容忍规则,确保Pod能被调度到带node=testvm:NoSchedule污点的节点:

apiVersion: v1
kind: Pod
metadata:
  name: nginx
spec:
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:
        - matchExpressions:
          - key: node
            operator: In
            values:
            - testvm 
  tolerations:
  - key: "node"
    operator: "Equal"
    value: "testvm"
    effect: "NoSchedule"
  containers:
  - name: nginx
    image: nginx
    imagePullPolicy: IfNotPresent

3. 重新创建Pod

执行创建命令:

kubectl apply -f <你的Pod清单文件名>.yaml

额外检查

如果仍出现亲和不匹配的警告,确认目标节点的Label是否正确设置:

kubectl get nodes --show-labels | grep testvm

确保节点确实带有node=testvm的Label。


内容的提问来源于stack exchange,提问作者Container-Man

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 15:20:21