如何让AKS优先扩容Spot实例节点池?配置问题排查
问题分析与解决方案
当前配置的问题
- 亲和性权重过低:Pod使用的
preferredDuringSchedulingIgnoredDuringExecution亲和性权重仅为1,调度器和Cluster Autoscaler会认为这种偏好优先级极低,不足以优先选择Spot节点池。 - 节点池无扩容器优先级设置:未给两个节点池配置Cluster Autoscaler的优先级注解,导致Autoscaler在评估扩容目标时没有明确的优先级导向,会默认选择无额外污点限制的常规节点池。
实现优先扩容Spot节点池的步骤
1. 给节点池设置Cluster Autoscaler优先级
通过注解给Spot节点池设置更高的优先级数值(数值越大优先级越高),常规节点池设置更低数值:
# 给gpuspot1节点池设置高优先级 kubectl annotate nodepool gpuspot1 cluster-autoscaler.kubernetes.io/node-pool-priority="10" --namespace=kube-system # 给gpuscale1节点池设置低优先级 kubectl annotate nodepool gpuscale1 cluster-autoscaler.kubernetes.io/node-pool-priority="1" --namespace=kube-system
2. 提升Pod的Spot节点偏好权重
将Pod亲和性中的权重调高,强化调度器对Spot节点的偏好:
tolerations: - key: sku value: gpu effect: NoSchedule - key: kubernetes.azure.com/scalesetpriority value: spot effect: "NoSchedule" affinity: nodeAffinity: preferredDuringSchedulingIgnoredDuringExecution: - weight: 100 preference: matchExpressions: - key: kubernetes.azure.com/scalesetpriority operator: In values: - spot
3. 配置生效逻辑
调整后,当有Pod需要调度时:
- 调度器会优先尝试将Pod调度到Spot节点(因亲和性权重高)
- 若Spot节点池无足够资源,Cluster Autoscaler会优先扩容Spot节点池(因优先级注解)
- 当Spot实例因Azure容量不足无法扩容时,Autoscaler会自动切换到常规节点池gpuscale1进行扩容,符合需求。
内容的提问来源于stack exchange,提问作者GEWLAR
相关产品推荐
相关产品推荐

