You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

EKS Fargate环境下CoreDNS Pod持续处于Pending状态求助

问题分析与解决方案

问题背景

在私有子网中部署的EKS集群(配置了Fargate配置文件),执行命令移除CoreDNS Deployment的eks.amazonaws.com/compute-type注解并重启后,CoreDNS Pod持续处于Pending状态,事件日志显示Pod provisioning timed out。

核心原因

  1. 调度目标缺失:移除compute-type=fargate注解后,Pod不再被Fargate调度,若集群未配置EC2节点组,Pod无可用节点可调度。
  2. 私有子网依赖缺失:私有子网内的Fargate Pod需要特定VPC端点才能完成镜像拉取、与EKS控制平面通信等操作,缺失则会导致调度超时。
  3. Fargate调度逻辑冲突:Pod仍带有Fargate相关标签,Fargate配置文件选择器匹配该Pod,但因网络问题无法完成资源创建。

解决方案1:恢复Fargate调度(继续用Fargate运行CoreDNS)

若原本计划用Fargate承载CoreDNS,无需移除该注解,直接恢复配置:

kubectl patch deployment coredns \
  -n kube-system \
  --type merge \
  -p '{"spec":{"template":{"metadata":{"annotations":{"eks.amazonaws.com/compute-type":"fargate"}}}}}'

重启Pod生效:

kubectl rollout restart deployment coredns -n kube-system

解决方案2:配置私有子网Fargate所需VPC端点

私有子网内的Fargate必须依赖以下VPC端点才能正常运行,添加后可解决调度超时问题:

  • ECR镜像拉取相关:ecr.dkr、ecr.api、S3网关端点
  • EKS控制平面通信:eks、eks-fargate

Terraform配置示例

# ECR DKR接口端点(镜像拉取)
resource "aws_vpc_endpoint" "ecr_dkr" {
  vpc_id            = var.vpc_id
  service_name      = "com.amazonaws.${var.region}.ecr.dkr"
  vpc_endpoint_type = "Interface"

  security_group_ids = [var.cluster_security_group_id]
  subnet_ids         = var.subnet_ids

  private_dns_enabled = true
}

# ECR API接口端点(镜像仓库API访问)
resource "aws_vpc_endpoint" "ecr_api" {
  vpc_id            = var.vpc_id
  service_name      = "com.amazonaws.${var.region}.ecr.api"
  vpc_endpoint_type = "Interface"

  security_group_ids = [var.cluster_security_group_id]
  subnet_ids         = var.subnet_ids

  private_dns_enabled = true
}

# S3网关端点(ECR镜像存储访问)
resource "aws_vpc_endpoint" "s3" {
  vpc_id            = var.vpc_id
  service_name      = "com.amazonaws.${var.region}.s3"
  vpc_endpoint_type = "Gateway"

  route_table_ids = var.private_subnet_route_table_ids
}

# EKS接口端点(控制平面通信)
resource "aws_vpc_endpoint" "eks" {
  vpc_id            = var.vpc_id
  service_name      = "com.amazonaws.${var.region}.eks"
  vpc_endpoint_type = "Interface"

  security_group_ids = [var.cluster_security_group_id]
  subnet_ids         = var.subnet_ids

  private_dns_enabled = true
}

# EKS Fargate接口端点(调度逻辑通信)
resource "aws_vpc_endpoint" "eks_fargate" {
  vpc_id            = var.vpc_id
  service_name      = "com.amazonaws.${var.region}.eks-fargate"
  vpc_endpoint_type = "Interface"

  security_group_ids = [var.cluster_security_group_id]
  subnet_ids         = var.subnet_ids

  private_dns_enabled = true
}

应用配置后,重启CoreDNS Pod即可。


解决方案3:添加EC2节点组(切换到EC2运行CoreDNS)

若计划用EC2节点承载CoreDNS,需创建EKS节点组:

Terraform配置示例

resource "aws_eks_node_group" "core_node_group" {
  cluster_name    = var.cluster_name
  node_group_name = "core-dns-node-group"
  node_role_arn   = var.node_iam_role_arn
  subnet_ids      = var.subnet_ids

  scaling_config {
    desired_size = 2
    max_size     = 3
    min_size     = 1
  }

  tags = var.tags
}

注意事项

  • 确保节点IAM角色拥有AmazonEKSWorkerNodePolicy、AmazonEC2ContainerRegistryReadOnly等必要权限
  • 等待节点加入集群后,CoreDNS Pod会自动调度到EC2节点

额外检查点

  • 确认Fargate配置文件的selector.namespace包含kube-system,确保Fargate能匹配该namespace的Pod
  • 检查安全组规则:允许Fargate安全组出站访问VPC端点及EKS控制平面
  • 查看Pod事件详情:kubectl describe pod <coredns-pod-name> -n kube-system,确认是否有其他网络或权限报错

内容的提问来源于stack exchange,提问作者JustARandomGuy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 07:35:36