You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GCP托管K8s集群Web应用Pod跨节点池分布配置异常问题

解决GKE中Web应用Pod全调度至Spot VM的问题

问题根源

你使用的preferredDuringSchedulingIgnoredDuringExecution是软调度偏好,当Spot节点池有可用资源时,Kubernetes会优先将所有符合条件的Pod调度过去,无法保证常规VM节点池保留指定数量的Pod。

解决方案

方案1:拆分Deployment精确控制Pod分布(推荐)

将Web应用拆分为两个独立的Deployment,分别控制常规VM和Spot VM上的Pod数量:

  1. 常规VM专用Deployment(固定1个Pod)
    配置强制节点亲和性,确保Pod只能调度到常规VM节点池:

    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: web-app-regular
    spec:
      replicas: 1
      selector:
        matchLabels:
          app: web-app
          tier: regular
      template:
        metadata:
          labels:
            app: web-app
            tier: regular
        spec:
          affinity:
            nodeAffinity:
              requiredDuringSchedulingIgnoredDuringExecution:
                nodeSelectorTerms:
                - matchExpressions:
                  - key: node-pool-type
                    operator: In
                    values:
                    - regular
          # 其余Pod模板配置(容器、资源等)与原应用一致
    
  2. Spot VM专用Deployment(2-4个Pod)
    配置偏好节点亲和性,优先调度到Spot VM节点池:

    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: web-app-spot
    spec:
      replicas: 3 # 根据需求调整为2-4
      selector:
        matchLabels:
          app: web-app
          tier: spot
      template:
        metadata:
          labels:
            app: web-app
            tier: spot
        spec:
          affinity:
            nodeAffinity:
              preferredDuringSchedulingIgnoredDuringExecution:
              - weight: 100
                preference:
                  matchExpressions:
                  - key: node-pool-type
                    operator: In
                    values:
                    - spot
          # 其余Pod模板配置(容器、资源等)与原应用一致
    

方案2:使用拓扑分布约束控制Pod分布

如果不想拆分Deployment,可通过Kubernetes拓扑分布约束,强制保证常规VM节点池至少有1个Pod:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: web-app
spec:
  replicas: 4 # 3-5之间调整
  selector:
    matchLabels:
      app: web-app
  template:
    metadata:
      labels:
        app: web-app
    spec:
      affinity:
        nodeAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
          - weight: 100
            preference:
              matchExpressions:
              - key: node-pool-type
                operator: In
                values:
                - spot
      topologySpreadConstraints:
      - maxSkew: 1
        topologyKey: node-pool-type
        whenUnsatisfiable: DoNotSchedule
        labelSelector:
          matchLabels:
            app: web-app
      # 其余Pod模板配置

该配置会让Pod在node-pool-type(常规/Spot)两个拓扑域中尽可能均匀分布,保证常规VM上至少有1个Pod(当总副本数≥1时)。

前置准备:确保节点池标签正确

先确认常规VM和Spot VM节点池已打上对应标签:

  • 给常规VM节点池打标签:
    gcloud container node-pools update YOUR_REGULAR_POOL_NAME --cluster YOUR_CLUSTER_NAME --update-labels node-pool-type=regular
    
  • 给Spot VM节点池打标签:
    gcloud container node-pools update YOUR_SPOT_POOL_NAME --cluster YOUR_CLUSTER_NAME --update-labels node-pool-type=spot
    

额外检查

  • 移除可能导致Pod聚集到Spot节点的podAffinity规则,或调整规则避免冲突。
  • 确保常规VM节点池有足够资源容纳至少1个Web应用Pod。

内容的提问来源于stack exchange,提问作者brian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 12:15:26