You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

EC2自动伸缩组中ES升级7+后cluster.initial_master_nodes配置问题

Elasticsearch 6.8 升级至 8.9:ASG 环境下 cluster.initial_master_nodes 配置问题

我们正在将AWS EC2实例上的Elasticsearch从6.8版本升级到8.9版本,核心卡在cluster.initial_master_nodes的配置上:

  • ES 8.x强制要求配置该参数,且需要指定固定的node.name列表,但实例由Terraform管理的自动伸缩组(ASG)创建,每个ASG对应唯一的ES模板,无法给每个实例传递索引编号生成固定名称。
  • 依赖ASG维持固定数量的实例,暂时没有替代方案。
  • 6.8版本无需配置cluster.initial_master_nodes,直接用${HOSTNAME}作为node.name即可;实例主机名格式为ip-${privateIP}.ec2.internal,但ASG自动创建的实例IP无法提前预知,没法预先配置IP列表。

可行解决方案

方案1:利用EC2元数据+API动态生成配置

通过实例元数据获取自身标识,结合EC2 API拉取同集群master节点列表,动态生成node.name和cluster.initial_master_nodes配置,无需提前知晓实例信息。

修改后的cloud-config配置

write_files:
  - path: "/elasticsearch/config/elasticsearch.yml"
    content: |
      cluster.name: ${cluster_name}
      cluster.routing.allocation.awareness.attributes: aws_availability_zone
      cloud.node.auto_attributes: true
      xpack.security.http.ssl.enabled: false
      xpack.security.enabled: false
      network.host: 0.0.0.0
      network.publish_host: _ec2_
      http.port: 9200
      http.compression: true
      transport.port: 9300
      discovery.seed_providers: ec2
      discovery.ec2.availability_zones: ${availability_zones}
      discovery.ec2.protocol: http
      discovery.ec2.host_type: private_ip
      discovery.ec2.endpoint: ${aws_ec2_endpoint}
      discovery.ec2.tag.ProductCode: ${product_code_tag}
      discovery.ec2.tag.role: ${master_node}
      discovery.ec2.tag.Environment: ${environment_tag}
      discovery.ec2.tag.InventoryCode: ${inventory_code_tag}
      logger.discovery: DEBUG
      node.name: ${cluster_name}-master-temp
      node.roles: ${node_roles}
      script.painless.regex.enabled: true
      thread_pool.search.size: 200
      thread_pool.search.queue_size: 20000
      thread_pool.write.size: ${bulk_threads_size}
      thread_pool.write.queue_size: 20000
  - path: "/opt/setup-es-cluster.sh"
    permissions: "0755"
    content: |
      #!/bin/bash
      # 获取实例ID后缀作为唯一标识
      INSTANCE_ID=$(curl -s http://169.254.169.254/latest/meta-data/instance-id)
      INSTANCE_SUFFIX=${INSTANCE_ID: -8}
      # 替换node.name为固定值
      sed -i "s/${cluster_name}-master-temp/${cluster_name}-master-${INSTANCE_SUFFIX}/g" /elasticsearch/config/elasticsearch.yml

      # 拉取同集群所有master节点的标识列表
      MASTER_NODE_LIST=$(aws ec2 describe-instances \
        --filters "Name=tag:role,Values=${master_node}" "Name=tag:Environment,Values=${environment_tag}" "Name=tag:InventoryCode,Values=${inventory_code_tag}" "Name=tag:ProductCode,Values=${product_code_tag}" \
        --query 'Reservations[].Instances[].InstanceId' \
        --output text | awk -v cluster=${cluster_name} '{print "\""cluster"-master-"substr($1, length($1)-7)"\""}' | tr '\n' ',')
      # 移除最后一个逗号并添加配置
      MASTER_NODE_LIST=${MASTER_NODE_LIST%,}
      echo "cluster.initial_master_nodes: [${MASTER_NODE_LIST}]" >> /elasticsearch/config/elasticsearch.yml

runcmd:
  - /opt/setup-es-cluster.sh

配套调整

确保实例绑定的IAM角色拥有ec2:DescribeInstances权限,Terraform原有配置无需额外修改。


方案2:固定节点名称+多Launch Configuration(适用于固定实例数的ASG)

如果ASG始终维持固定数量的实例(比如3个),通过Terraform的count为每个实例生成唯一索引,创建对应Launch Configuration,直接配置固定的node.name和cluster.initial_master_nodes列表。

修改后的Terraform配置

data "template_file" "master-es" {
  count = var.master_nodes_count
  template = file("es-cloud-config.yaml")
  vars = {
    cluster_name            = var.cluster_name
    heap_size               = var.master_heap_size
    availability_zones      = join(",", data.aws_availability_zones.available.names)
    min_master_nodes        = module.min-master-nodes.value
    master_node             = "true"
    data_node               = "false"
    docker_elasticsearch_tag= var.docker_elasticsearch_tag
    bulk_threads_size       = var.master_bulk_threads_size
    node_index              = count.index
    master_node_count       = var.master_nodes_count
  }
}

resource "aws_launch_configuration" "master-es" {
  count                = var.master_nodes_count
  name_prefix          = "my-master-${count.index}"
  image_id             = module.ami.ami_id
  instance_type        = var.master_instance_type
  security_groups      = var.my_security_groups
  user_data            = data.template_file.master-es[count.index].rendered
  key_name             = var.ssh_key
  iam_instance_profile = aws_iam_instance_profile.default.name
  enable_monitoring    = true

  root_block_device {
    volume_size = var.master_storage["volume_size"]
    volume_type = var.master_storage["volume_type"]
    iops = var.master_storage["volume_iops"]
    delete_on_termination = "true"
  }

  lifecycle {
    create_before_destroy = true
  }
}

resource "aws_autoscaling_group" "master-es" {
  count                     = var.master_nodes_count > 0 ? 1 : 0
  name                      = "my-master-asg"
  min_size                  = var.master_nodes_count
  max_size                  = var.master_nodes_count
  desired_capacity          = var.master_nodes_count
  force_delete              = true
  launch_configurations     = [for idx in range(var.master_nodes_count) : aws_launch_configuration.master-es[idx].id]
  load_balancers            = [aws_elb.master-es[count.index].id]
  vpc_zone_identifier       = var.my_subnets
  termination_policies      = ["OldestInstance"]
  wait_for_capacity_timeout = var.master_capacity_timeout
  min_elb_capacity          = var.master_nodes_count
  health_check_type         = "EC2"

  lifecycle {
    create_before_destroy = true
  }
}

修改后的cloud-config配置

write_files:
  - path: "/elasticsearch/config/elasticsearch.yml"
    content: |
      cluster.name: ${cluster_name}
      cluster.routing.allocation.awareness.attributes: aws_availability_zone
      cloud.node.auto_attributes: true
      xpack.security.http.ssl.enabled: false
      xpack.security.enabled: false
      network.host: 0.0.0.0
      network.publish_host: _ec2_
      http.port: 9200
      http.compression: true
      transport.port: 9300
      discovery.seed_providers: ec2
      discovery.ec2.availability_zones: ${availability_zones}
      discovery.ec2.protocol: http
      discovery.ec2.host_type: private_ip
      discovery.ec2.endpoint: ${aws_ec2_endpoint}
      discovery.ec2.tag.ProductCode: ${product_code_tag}
      discovery.ec2.tag.role: ${master_node}
      discovery.ec2.tag.Environment: ${environment_tag}
      discovery.ec2.tag.InventoryCode: ${inventory_code_tag}
      logger.discovery: DEBUG
      node.name: ${cluster_name}-master-${node_index}
      cluster.initial_master_nodes: [${join(",", formatlist("\"${cluster_name}-master-%d\"", range(master_node_count)))}]
      node.roles: ${node_roles}
      script.painless.regex.enabled: true
      thread_pool.search.size: 200
      thread_pool.search.queue_size: 20000
      thread_pool.write.size: ${bulk_threads_size}
      thread_pool.write.queue_size: 20000

方案3:跳过初始化配置(仅适用于已初始化集群的滚动升级)

如果是在已有运行集群的基础上滚动升级,可以先通过临时节点完成ES 8.x集群初始化,之后移除cluster.initial_master_nodes参数,ASG节点通过EC2发现自动加入集群。

操作步骤

  1. 手动启动一个ES 8.x临时master节点,配置好cluster.initial_master_nodes完成集群初始化。
  2. 移除ASG模板中的cluster.initial_master_nodes配置,启动ASG节点,通过EC2发现加入集群。
  3. 逐步替换旧的6.8版本节点。

内容的提问来源于stack exchange,提问作者Zefr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 14:54:58