如何配置AKS集群:让Spot实例优先调度Pod,终止后切换至按需节点池
实现AKS集群Spot节点优先调度的方法
针对你的需求——让Spot实例优先承接Pod调度,仅当Spot节点被终止后才使用按需节点池,结合AKS中Spot节点默认的污点(kubernetes.azure.com/scalesetpriority=spot:NoSchedule)和标签(kubernetes.azure.com/scalesetpriority=spot),可以通过以下几种方式实现:
方法一:污点容忍+强制节点亲和性+Pod优先级类
这是最直接的配置方式,通过三重规则确保Pod优先被调度到Spot节点:
- 添加污点容忍:给Pod配置容忍Spot节点的默认污点,让Pod能够被调度到带有该污点的节点上。
- 强制节点亲和性:使用
requiredDuringSchedulingIgnoredDuringExecution类型的节点亲和性,强制Pod在调度时优先选择带有Spot标签的节点;当Spot节点全部被回收后,亲和性条件无法满足,调度器会自动转向按需节点。 - 高优先级类:创建高优先级的Pod优先级类,确保这类Pod在调度队列中拥有最高优先级,优先抢占Spot节点资源。
示例配置
1. 创建高优先级类
apiVersion: scheduling.k8s.io/v1 kind: PriorityClass metadata: name: spot-priority value: 1000000 globalDefault: false description: "最高优先级,用于优先调度到Spot节点的Pod"
2. 带亲和性和容忍的Deployment配置
apiVersion: apps/v1 kind: Deployment metadata: name: spot-preferred-app spec: replicas: 5 selector: matchLabels: app: spot-app template: metadata: labels: app: spot-app spec: priorityClassName: spot-priority tolerations: - key: "kubernetes.azure.com/scalesetpriority" operator: "Equal" value: "spot" effect: "NoSchedule" affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: kubernetes.azure.com/scalesetpriority operator: In values: - spot containers: - name: nginx image: nginx:latest
方法二:Cluster Autoscaler节点优先级配置
通过调整Cluster Autoscaler的节点池扩容优先级,让Spot节点池优先扩容,结合Pod的污点容忍实现优先调度:
- 给Spot节点池添加高优先级的annotation,按需节点池设置较低优先级。
- 配置Cluster Autoscaler使用
priority扩容策略,确保需要扩容时优先选择Spot节点池。
配置步骤
- 更新Spot节点池的annotation:
az aks nodepool update --name spot-nodepool --cluster-name your-aks-cluster --set tags."cluster-autoscaler\.kubernetes\.io/priority"=10
- 更新按需节点池的annotation(设置更低优先级,比如1):
az aks nodepool update --name ondemand-nodepool --cluster-name your-aks-cluster --set tags."cluster-autoscaler\.kubernetes\.io/priority"=1
- 确保Cluster Autoscaler启用
priority扩容策略(AKS默认可能已配置,若未配置可通过AKS集群更新开启)。
方法三:自定义调度器打分规则
通过修改Kubernetes调度器的节点打分配置,让Spot节点获得更高的调度打分权重,确保调度器优先选择Spot节点:
- 自定义调度器的
schedulingProfile,调整nodeScorePlugins中的权重,比如提高NodeAffinity插件的权重,或者添加自定义打分逻辑。 - 配合Pod的污点容忍,确保Pod能够被调度到Spot节点。
这种方式复杂度较高,适合需要深度定制调度逻辑的场景,一般推荐前两种方法。
注意事项
- 所有需要优先调度到Spot节点的Pod,都需要配置对应的污点容忍和亲和性(或通过模板批量配置)。
- 当Spot节点被Azure回收时,Pod会被自动驱逐,调度器会将Pod重新调度到可用的按需节点上。
- 对于有状态应用,建议配置
PodDisruptionBudget,避免节点回收时导致业务中断。
内容的提问来源于stack exchange,提问作者sysadmincrispy
相关产品推荐
相关产品推荐

