You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

EKS Kubernetes升级至1.29后MongoDB Pod存储卷挂载失败求助

问题分析与解决方案

核心问题

AWS EKS Kubernetes版本升级至1.29后,MongoDB Pod因EBS卷挂载失败陷入ContainerCreating状态,其余服务Pod运行正常。错误根源为STS AssumeRoleWithWebIdentity 请求重试配额耗尽,导致无法获取凭证完成EC2实例信息查询与卷挂载操作。

排查与解决步骤

1. 校验EBS CSI驱动的IAM身份配置

  • 确认EBS CSI控制器ServiceAccount关联的IAM角色具备以下权限:
    • ec2:DescribeInstances
    • ec2:AttachVolume
    • ec2:DescribeVolumes
  • 检查IAM角色的信任策略,确保包含EKS OIDC提供商及对应ServiceAccount的身份条件,示例信任策略片段:
    {
      "Effect": "Allow",
      "Principal": {
        "Federated": "arn:aws:iam::ACCOUNT_ID:oidc-provider/oidc.eks.REGION.amazonaws.com/id/OIDC_ID"
      },
      "Action": "sts:AssumeRoleWithWebIdentity",
      "Condition": {
        "StringEquals": {
          "oidc.eks.REGION.amazonaws.com/id/OIDC_ID:sub": "system:serviceaccount:kube-system:ebs-csi-controller-sa"
        }
      }
    }
    

2. 调整STS API速率限制

  • 查看当前AWS账户STS区域速率配额,若已达上限,提交AWS支持工单申请提升配额;
  • 调整EBS CSI控制器的重试参数,减少频繁请求导致的配额耗尽:
    修改ebs-csi-controller Deployment的容器参数,添加--retry-max=3降低重试次数,或调整--retry-delay=5s增加重试间隔。

3. 验证节点IAM角色权限

  • 确认EKS节点的IAM角色包含sts:AssumeRoleWithWebIdentity权限,且未被IAM策略限制STS请求次数。

4. 临时缓解操作

  • 重启EBS CSI控制器Pod,重置凭证缓存:
    kubectl rollout restart deployment ebs-csi-controller -n kube-system
    
  • 删除故障MongoDB Pod,触发重新调度:
    kubectl delete pod mongo-6749cf47f6-54ddw mongo-78764f686f-vgvtg
    

原始问题信息

Pod状态详情

NAME                           READY   STATUS              RESTARTS   AGE
api-gateway-768d8b8cc6-f7swz   1/1     Running             0          8h
authapi-cfb964777-gnjvz        1/1     Running             0          8h
basketapi-7cbb86fc75-2xq2h     1/1     Running             0          8h
mongo-6749cf47f6-54ddw         0/1     ContainerCreating   0          4h17m
mongo-78764f686f-vgvtg         0/1     ContainerCreating   0          36m
packageapi-56854f7dbf-zqn4f    1/1     Running             0          8h
providerapi-7c78ddb486-m9bcf   1/1     Running             0          8h
serviceapi-6877d9b6b4-t4ggz    1/1     Running             0          8h
userapi-f5b8c64f4-gw565        1/1     Running             0          8h

完整错误信息

AttachVolume.Attach failed for volume "pvc-" : rpc error: code = Internal desc = Could not attach volume "vol-" to node "i-": error listing AWS instances: operation error EC2: DescribeInstances, get identity: get credentials: failed to refresh cached credentials, failed to retrieve credentials, operation error STS: AssumeRoleWithWebIdentity, failed to get rate limit token, retry quota exceeded, 0 available, 5 requested

详细事件信息

kubectl get event  
LAST SEEN   TYPE      REASON               OBJECT                       MESSAGE
92s         Warning   FailedAttachVolume   pod/mongo-6749cf47f6-54ddw   AttachVolume.Attach failed for volume "pvc-38e6ea4c-6d63-46a8-8b0f-0ccc74c05e8c" : rpc error: code = Internal desc = Could not attach volume "vol-08a4fca2489e3cef8" to node "i-09b9b0fde432495b9": error listing AWS instances: operation error EC2: DescribeInstances, get identity: get credentials: failed to refresh cached credentials, failed to retrieve credentials, operation error STS: AssumeRoleWithWebIdentity, failed to get rate limit token, retry quota exceeded, 0 available, 5 requested
92s         Warning   FailedAttachVolume   pod/mongo-78764f686f-vgvtg   AttachVolume.Attach failed for volume "pvc-38e6ea4c-6d63-46a8-8b0f-0ccc74c05e8c" : rpc error: code = Internal desc = Could not attach volume "vol-08a4fca2489e3cef8" to node "i-09b9b0fde432495b9": error listing AWS instances: operation error EC2: DescribeInstances, get identity: get credentials: failed to refresh cached credentials, failed to retrieve credentials, operation error STS: AssumeRoleWithWebIdentity, failed to get rate limit token, retry quota exceeded, 0 available, 5 requested

内容的提问来源于stack exchange,提问作者Pedro Paulino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 09:23:14