启用Pod反亲和性后更新Deployment时Pod调度失败的解决建议
问题分析与解决方案
问题根源
调度失败的核心是两条亲和性规则的冲突:
- nodeAffinity为强制要求:新Pod必须调度到
kubernetes.io/hostname=example.com节点 - podAntiAffinity为强制要求:新Pod不能与
component=myapp的Pod处于同一主机
当前example.com节点已运行旧的myapp Pod,新Pod既无法违反nodeAffinity去其他节点,也无法违反podAntiAffinity留在该节点,最终导致无可用调度节点。
具体解决建议
方案1:放宽nodeAffinity的节点范围
既然集群有3个节点,扩大nodeAffinity允许的节点列表,让Pod能调度到多个节点,配合反亲和性实现新旧Pod分散部署:
affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: kubernetes.io/hostname operator: In values: - example.com - node2.example.com # 补充其他节点的hostname - node3.example.com podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchExpressions: - key: component operator: In values: - myapp topologyKey: "kubernetes.io/hostname"
方案2:修正podAntiAffinity的偏好性配置
你之前尝试preferredDuringSchedulingIgnoredDuringExecution无效,大概率是配置格式错误。正确的偏好性反亲和性配置如下,当无其他节点可选时,调度器会允许新Pod与旧Pod同主机:
affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: kubernetes.io/hostname operator: In values: - example.com podAntiAffinity: preferredDuringSchedulingIgnoredDuringExecution: - weight: 100 # 权重越高,调度器越优先满足该规则 podAffinityTerm: labelSelector: matchExpressions: - key: component operator: In values: - myapp topologyKey: "kubernetes.io/hostname"
方案3:调整Deployment滚动更新策略
若必须保留原有亲和性规则,可修改滚动更新参数,让旧Pod先删除再创建新Pod,释放节点的反亲和性限制:
strategy: rollingUpdate: maxSurge: 0 # 不额外创建新Pod maxUnavailable: 1 # 每次删除一个旧Pod type: RollingUpdate
方案4:验证节点标签正确性
确认另外两个节点的kubernetes.io/hostname标签无拼写错误,确保nodeAffinity的节点列表与集群实际节点匹配:
kubectl get nodes --show-labels | grep kubernetes.io/hostname
内容的提问来源于stack exchange,提问作者Moin Memon
相关产品推荐
相关产品推荐

