You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

EKS 1.24+环境下双副本Elasticsearch STS卡在ContainerCreating状态

问题描述

将EKS开发环境升级至1.24版本后,Elasticsearch 7.17.5的StatefulSet实例卡在ContainerCreating状态,无任何事件或日志输出。经排查发现,仅当Elasticsearch集群配置2个副本时才会出现该问题,PVC、PV、ConfigMap、Secret等其他资源均正常,该配置在Kubernetes 1.23及以下版本、AKS环境中均可正常运行。偶尔强制删除Pod后能启动,但会提示节点上已有其他Elasticsearch实例运行。

Pod描述信息如下:

Name:                 elasticsearch-0
Namespace:            XXXXXXXXXXX
Priority:             2000000
Priority Class Name:  apps
Service Account:      es
Node:                 XXXXXXXXXXX
Start Time:           Fri, 30 Jun 2023 04:17:55 +0300
Labels:               app=elasticsearch
                      controller-revision-hash=elasticsearch-68574b48b5
                      owner=control-plane
                      statefulset.kubernetes.io/pod-name=elasticsearch-0
Annotations:          <none>
Status:               Pending
IP:                    
IPs:                  <none>
Controlled By:        StatefulSet/elasticsearch
Containers:
  elastic:
    Container ID:   
    Image:          es:7.17.5
    Image ID:       
    Ports:          9200/TCP, 9300/TCP
    Host Ports:     0/TCP, 0/TCP
    State:          Waiting
      Reason:       ContainerCreating
    Ready:          False
    Restart Count:  0
    Limits:
      cpu:     4
      memory:  8Gi
    Requests:
      cpu:      500m
      memory:   4Gi
    Readiness:  exec [bash -c set -e
# If the node is starting up wait for the cluster to be ready (request params: "wait_for_status=yellow&timeout=1s" )
# Once it has started only check that the node itself is responding
START_FILE=/tmp/.es_start_file

# Disable nss cache to avoid filling dentry cache when calling curl
# This is required with Elasticsearch Docker using nss < 3.52
export NSS_SDB_USE_CACHE=no

http () {
  local path="${1}"
  local args="${2}"
  set -- -XGET -s

  if [ "$args" != "" ]; then
    set -- "$@" $args
  fi

  if [ -n "${ES_PASS}" ]; then
    set -- "$@" -u "${ES_USER}:${ES_PASS}"
  fi

  curl --output /dev/null -k "$@" "http://127.0.0.1:9200${path}"
}

if [ -f "${START_FILE}" ]; then
  echo 'Elasticsearch is already running, lets check the node is healthy'
  HTTP_CODE=$(http "/" "-w %{http_code}")
  RC=$?
  if [[ ${RC} -ne 0 ]]; then
    echo "curl --output /dev/null -k -XGET -s -w '%{http_code}' \${BASIC_AUTH} http://127.0.0.1:9200/ failed with RC ${RC}"
    exit ${RC}
  fi
  # ready if HTTP code 200, 503 is tolerable if ES version is 6.x
  if [[ ${HTTP_CODE} == "200" ]]; then
    exit 0
  elif [[ ${HTTP_CODE} == "503" && "7" == "6" ]]; then
    exit 0
  else
    echo "curl --output /dev/null -k -XGET -s -w '%{http_code}' \${BASIC_AUTH} http://127.0.0.1:9200/ failed with HTTP code ${HTTP_CODE}"
    exit 1
  fi

else
  echo 'Waiting for elasticsearch cluster to become ready (request params: "wait_for_status=yellow&timeout=1s" )'
  if http "/_cluster/health?wait_for_status=yellow&timeout=1s" "--fail" ; then
    touch ${START_FILE}
    exit 0
  else
    echo 'Cluster is not yet ready (request params: "wait_for_status=yellow&timeout=1s" )'
    exit 1
  fi
fi
] delay=10s timeout=5s period=10s #success=3 #failure=3
    Environment Variables from:
      es-creds   Secret     Optional: false
      es-ilm-cm  ConfigMap  Optional: false
    Environment:
      cluster.name:                                          es
      node.name:                                             elasticsearch-0 (v1:metadata.name)
      ELASTIC_PASSWORD:                                      <set to the key 'ES_PASS' in secret 'es-creds'>  Optional: false
      ingest.geoip.downloader.enabled:                       false
      network.host:                                          0.0.0.0
      cluster.initial_master_nodes:                          elasticsearch-0,elasticsearch-1,
      discovery.seed_hosts:                                  elasticsearch-headless
      cluster.deprecation_indexing.enabled:                  false
      path.data:                                             /usr/share/elasticsearch/data/data
      path.logs:                                             /usr/share/elasticsearch/data/logs
      ES_JAVA_OPTS:                                          -Xms2g -Xmx2g
      xpack.security.enabled:                                true
      xpack.security.authc.realms.native.native1.order:      0
      xpack.security.authc.realms.file.file1.order:          1
      xpack.security.transport.ssl.enabled:                  true
      xpack.security.transport.ssl.verification_mode:        certificate
      xpack.security.transport.ssl.key:                      /usr/share/elasticsearch/config/certs/tls.key
      xpack.security.transport.ssl.certificate:              /usr/share/elasticsearch/config/certs/tls.crt
      xpack.security.transport.ssl.certificate_authorities:  /usr/share/elasticsearch/config/certs/ca.crt
    Mounts:
      /tmp/elastic/ from es-ilm (rw)
      /usr/share/elasticsearch/config/certs from es-certs (ro)
      /usr/share/elasticsearch/data from es-storage (rw)
      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-rrwh9 (ro)
Conditions:
  Type              Status
  Initialized       True 
  Ready             False 
  ContainersReady   False 
  PodScheduled      True 
Volumes:
  es-storage:
    Type:       PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
    ClaimName:  es-storage-elasticsearch-0
    ReadOnly:   false
  es-ilm:
    Type:      ConfigMap (a volume populated by a ConfigMap)
    Name:      es-ilm
    Optional:  false
  es-certs:
    Type:        Secret (a volume populated by a Secret)
    SecretName:  es-certs
    Optional:    false
  kube-api-access-rrwh9:
    Type:                    Projected (a volume that contains injected data from multiple sources)
    TokenExpirationSeconds:  3607
    ConfigMapName:           kube-root-ca.crt
    ConfigMapOptional:       <nil>
    DownwardAPI:             true
QoS Class:                   Burstable
Node-Selectors:              <none>
Tolerations:                 node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
                             node.kubernetes.io/unreachable:NoExecute op=Exists for 300s
Events:                      <none>

可能的原因分析

  • Pod同节点调度引发资源冲突:你的StatefulSet未配置Pod反亲和性,2个副本可能被调度到同一节点。Elasticsearch运行时会占用特定内存映射区域或临时资源,同节点部署时易触发冲突,导致容器无法创建。你提到删Pod后启动提示节点有其他实例,也验证了同节点调度的可能性。建议添加强制反亲和规则,确保Elasticsearch Pod分散到不同节点:

    affinity:
      podAntiAffinity:
        requiredDuringSchedulingIgnoredDuringExecution:
        - labelSelector:
            matchExpressions:
            - key: app
              operator: In
              values:
              - elasticsearch
          topologyKey: kubernetes.io/hostname
    
  • EKS 1.24容器运行时切换的适配问题:EKS 1.24默认使用containerd替代dockershim作为容器运行时,部分Elasticsearch镜像的启动逻辑在containerd环境下存在并发创建容器的兼容性问题。当部署2个副本时,并发创建容器触发了该问题,而1个副本无并发压力、3个副本调度分散则不会触发。可以尝试拉取官方标准镜像测试,或检查镜像内的启动脚本是否有依赖docker特性的逻辑。

  • PersistentVolume并发挂载限制:EKS使用的CSI插件(如EBS CSI)可能对同一节点的并发PV挂载存在限制。当部署2个副本且被调度到同一节点时,同时挂载两个EBS卷可能触发插件的并发限制,导致容器创建卡住。1个副本无并发挂载需求,3个副本分散到不同节点也不会触发该限制。

  • 初始主节点配置的语法与集群quorum冲突:你的cluster.initial_master_nodes配置末尾存在多余逗号(elasticsearch-0,elasticsearch-1,),可能导致Elasticsearch解析时认为存在第三个未部署的主节点。当部署2个副本时,集群会一直等待第三个节点加入,虽不会直接导致ContainerCreating,但可能间接影响容器启动后的探针逻辑,进而被StatefulSet的调度逻辑误判,引发创建卡住的现象。

内容的提问来源于stack exchange,提问作者Israel Murciano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 23:56:58