You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes集群Pod副本跨节点调度失败问题排查

EKS集群Pod调度问题分析与解决

问题背景

集群包含3个节点:1个带污点的OnDemand节点、2个Spot节点。需求为:部署2个Pod副本,其中一个在OnDemand节点,另一个优先部署在Spot节点;同时要求相同Pod不能部署在同一节点。

已配置的Deployment片段:

tolerations:
  - key: dev-group
    value: OnDemand
    effect: NoSchedule

affinity:
  nodeAffinity:
    requiredDuringSchedulingIgnoredDuringExecution:
      nodeSelectorTerms:
      - matchExpressions:
        - key: eks.amazonaws.com/capacityType
          operator: In
          values:
          - ON_DEMAND
    preferredDuringSchedulingIgnoredDuringExecution:
    - weight: 1
      preference:
        matchExpressions:
        - key: eks.amazonaws.com/capacityType
          operator: In
          values:
          - SPOT
  podAntiAffinity:
    requiredDuringSchedulingIgnoredDuringExecution:
      - labelSelector:
          matchExpressions:
            - key: app.kubernetes.io/name
              operator: In
              values:
              - [app-name]
        topologyKey: kubernetes.io/hostname

预期效果:

  • 通过容忍度允许Pod调度至OnDemand节点
  • 要求Pod必须调度至ON_DEMAND节点
  • 优先将Pod调度至SPOT节点
  • 要求相同Pod不能部署在同一节点

实际调度错误:

0/3 nodes are available: 1 node(s) didn't match pod anti-affinity rules, 2 node(s) didn't match Pod's node affinity/selector. preemption: 0/3 nodes are available: 1 No preemption victims found for incoming pod, 2 Preemption is not helpful for scheduling..

已确认节点标签正确,移除podAntiAffinity后Pod可正常调度至OnDemand节点。

问题根源

核心矛盾在于nodeAffinity的强制约束与需求冲突:

  1. requiredDuringSchedulingIgnoredDuringExecution配置强制要求Pod只能调度到ON_DEMAND类型节点,直接排除了2个Spot节点;
  2. 仅剩1个OnDemand节点,但副本数为2,加上podAntiAffinity要求Pod不能同节点部署,没有足够节点满足调度条件,因此报错。

同时你的预期存在逻辑矛盾:既要求Pod必须部署在OnDemand节点,又希望其中一个副本部署在Spot节点,这两个要求无法同时满足。

修正方案

调整配置以匹配实际需求:允许Pod调度到两种类型节点,优先选择Spot,同时确保至少一个副本部署在OnDemand节点,且两个副本不同节点。

修正后的Deployment配置片段:

tolerations:
  - key: dev-group
    value: OnDemand
    effect: NoSchedule

affinity:
  nodeAffinity:
    # 允许Pod调度到ON_DEMAND或SPOT节点
    requiredDuringSchedulingIgnoredDuringExecution:
      nodeSelectorTerms:
      - matchExpressions:
        - key: eks.amazonaws.com/capacityType
          operator: In
          values:
          - ON_DEMAND
          - SPOT
    # 调高权重增强Spot节点的优先性
    preferredDuringSchedulingIgnoredDuringExecution:
    - weight: 100
      preference:
        matchExpressions:
        - key: eks.amazonaws.com/capacityType
          operator: In
          values:
          - SPOT
  # 强制Pod不同节点部署
  podAntiAffinity:
    requiredDuringSchedulingIgnoredDuringExecution:
      - labelSelector:
          matchExpressions:
            - key: app.kubernetes.io/name
              operator: In
              values:
              - [app-name]
        topologyKey: kubernetes.io/hostname

# 确保两种节点类型各至少有1个副本
topologySpreadConstraints:
- maxSkew: 1
  topologyKey: eks.amazonaws.com/capacityType
  whenUnsatisfiable: DoNotSchedule
  labelSelector:
    matchExpressions:
      - key: app.kubernetes.io/name
        operator: In
        values:
        - [app-name]

配置说明:

  • 修改nodeAffinity的required规则,允许Pod调度到两种节点类型,解决Spot节点被排除的问题;
  • 调高Spot节点的优先权重,确保调度器优先选择Spot节点;
  • 添加topologySpreadConstraints,强制Pod在ON_DEMAND和SPOT两种节点类型中各至少部署1个副本,满足"一个在OnDemand、一个在Spot"的核心需求;
  • 保留原有的podAntiAffinity和tolerations规则,确保Pod不同节点部署且能调度到带污点的OnDemand节点。

内容的提问来源于stack exchange,提问作者prosto.vint

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 02:33:23