AWS EKS节点组升级遇PodEvictionFailure错误的解决方法
使用CDK升级AWS EKS节点组时触发以下错误:
Resource handler returned message: "[ErrorDetail(ErrorCode=PodEvictionFailure, ErrorMessage=Reached max retries while trying to evict pods from nodes in node group <node-group-name>, ResourceIds=[<node-name>])] (Service: null, Status Code: 0, Request ID: null)" (RequestToken: <request-token>, HandlerErrorCode: GeneralServiceException)
根据AWS官方逻辑:
Deployment tolerating all the taints – Once every pod is evicted, it's expected for the node to be empty because the node is tainted in the earlier steps. However, if the deployment tolerates every taint, then the node is more likely to be non-empty, leading to pod eviction failure.
排查后发现,kube-system命名空间下的aws-node、kube-proxy两类DaemonSet Pod包含全局容忍规则({"operator": "Exists"}),具体配置如下:
{ "tolerations": [ { "operator": "Exists" }, { "key": "node.kubernetes.io/not-ready", "operator": "Exists", "effect": "NoExecute" }, { "key": "node.kubernetes.io/unreachable", "operator": "Exists", "effect": "NoExecute" }, { "key": "node.kubernetes.io/disk-pressure", "operator": "Exists", "effect": "NoSchedule" }, { "key": "node.kubernetes.io/memory-pressure", "operator": "Exists", "effect": "NoSchedule" }, { "key": "node.kubernetes.io/pid-pressure", "operator": "Exists", "effect": "NoSchedule" }, { "key": "node.kubernetes.io/unschedulable", "operator": "Exists", "effect": "NoSchedule" }, { "key": "node.kubernetes.io/network-unavailable", "operator": "Exists", "effect": "NoSchedule" } ] }
需解决以下问题:
- 是否需要修改这些AWS托管Pod的容忍配置?
- 具体操作方案是什么?
- 还有哪些手段可以避免PodEvictionFailure错误?
1. 不建议修改托管Pod的容忍配置
aws-node(负责VPC CNI网络)和kube-proxy是EKS集群的核心托管组件,AWS默认的全局容忍规则是为了确保它们能在各类节点状态下稳定运行。强行修改可能导致组件异常、网络中断甚至集群不可用,因此禁止直接修改这些Pod的原生容忍配置。
2. 针对本次升级的快速修复方案
开启节点组强制更新
在CDK的节点组配置中启用forceUpdateEnabled参数,当节点无法正常清空Pod时,AWS会直接终止节点,跳过驱逐重试流程。该方案适合能容忍短暂Pod中断的业务场景。
CDK代码示例(TypeScript):
import { Nodegroup } from 'aws-cdk-lib/aws-eks'; const nodeGroup = new Nodegroup(this, 'TargetNodeGroup', { cluster: myEksCluster, // 其他基础配置(如实例类型、子网等) forceUpdateEnabled: true, // 开启强制更新 });
手动提前清空节点
如果不想开启强制更新,可在触发CDK升级前手动清空目标节点:
# 忽略DaemonSet Pod(aws-node、kube-proxy属于此类),删除临时存储类Pod kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data
完成清空后再执行CDK部署,即可避免驱逐失败。
3. 长期规避PodEvictionFailure的通用措施
- 检查自定义Pod容忍规则:排查业务Pod的配置,避免使用
operator: Exists这类全局容忍规则,防止Pod滞留在待升级节点。 - 配置Pod中断预算(PDB):为核心业务Pod设置合理的PDB,平衡高可用与升级时的驱逐灵活性,避免因PDB限制导致驱逐失败。
- 调整升级批次策略:在CDK节点组配置中,通过
updateConfig参数控制单次升级的节点数量:const nodeGroup = new Nodegroup(this, 'TargetNodeGroup', { cluster: myEksCluster, updateConfig: { maxUnavailable: 1, // 单次升级最多不可用1个节点 // 或使用maxSurge: 1,升级时新增1个节点再替换旧节点 }, }); - 定期清理节点资源:升级前检查节点上的异常Pod(如无主Pod、Pending状态Pod),手动清理后再执行升级。
内容的提问来源于stack exchange,提问作者lipeiran

