Magento 2.4.3站点Elasticsearch偶发崩溃、CPU内存飙高问题求助
Elasticsearch 崩溃与资源占用异常问题优化记录
初始问题说明
我们的Elasticsearch偶尔出现崩溃,同时时常触发内存、CPU占用飙升,最终导致服务器无响应。初期我们保留了大部分默认配置,仅将JVM堆内存调整到48GB试图降低崩溃频率,后续调研发现Elasticsearch官方建议堆内存最大值不超过32GB,因此计划调整该参数。
服务器配置
CentOS 7 RAM: 125GB CPU: 40线程 硬盘: 2x Raid 1 NVME
硬件配置足以支撑现有业务,但处理现有数据需要进一步优化配置。当前我们运行的是Magento 2.4.3 CE版本站点,约有40万件商品。
初始配置文件
jvm.options文件
## JVM configuration ################################################################ ## IMPORTANT: JVM heap size ################################################################ ## ## You should always set the min and max JVM heap ## size to the same value. For example, to set ## the heap to 4 GB, set: ## ## -Xms4g ## -Xmx4g ## ## See https://www.elastic.co/guide/en/elasticsearch/reference/current/heap-size.html ## for more information ## ################################################################ # Xms represents the initial size of total heap space # Xmx represents the maximum size of total heap space -Xms48g -Xmx48g ################################################################ ## Expert settings ################################################################ ## ## All settings below this section are considered ## expert settings. Don't tamper with them unless ## you understand what you are doing ## ################################################################ ## GC configuration 8-13:-XX:+UseConcMarkSweepGC 8-13:-XX:CMSInitiatingOccupancyFraction=75 8-13:-XX:+UseCMSInitiatingOccupancyOnly ## G1GC Configuration # NOTE: G1 GC is only supported on JDK version 10 or later # to use G1GC, uncomment the next two lines and update the version on the # following three lines to your version of the JDK # 10-13:-XX:-UseConcMarkSweepGC # 10-13:-XX:-UseCMSInitiatingOccupancyOnly 14-:-XX:+UseG1GC 14-:-XX:G1ReservePercent=25 14-:-XX:InitiatingHeapOccupancyPercent=30 ## DNS cache policy # cache ttl in seconds for positive DNS lookups noting that this overrides the # JDK security property networkaddress.cache.ttl; set to -1 to cache forever -Des.networkaddress.cache.ttl=60 # cache ttl in seconds for negative DNS lookups noting that this overrides the # JDK security property networkaddress.cache.negative ttl; set to -1 to cache # forever -Des.networkaddress.cache.negative.ttl=10 ## optimizations # pre-touch memory pages used by the JVM during initialization -XX:+AlwaysPreTouch ## basic # explicitly set the stack size -Xss1m # set to headless, just in case -Djava.awt.headless=true # ensure UTF-8 encoding by default (e.g. filenames) -Dfile.encoding=UTF-8 # use our provided JNA always versus the system one -Djna.nosys=true # turn off a JDK optimization that throws away stack traces for common # exceptions because stack traces are important for debugging -XX:-OmitStackTraceInFastThrow # enable helpful NullPointerExceptions (https://openjdk.java.net/jeps/358), if # they are supported 14-:-XX:+ShowCodeDetailsInExceptionMessages # flags to configure Netty -Dio.netty.noUnsafe=true -Dio.netty.noKeySetOptimization=true -Dio.netty.recycler.maxCapacityPerThread=0 # log4j 2 -Dlog4j.shutdownHookEnabled=false -Dlog4j2.disable.jmx=true -Djava.io.tmpdir=${ES_TMPDIR} ## heap dumps # generate a heap dump when an allocation from the Java heap fails # heap dumps are created in the working directory of the JVM -XX:+HeapDumpOnOutOfMemoryError # specify an alternative path for heap dumps; ensure the directory exists and # has sufficient space -XX:HeapDumpPath=/var/lib/elasticsearch # specify an alternative path for JVM fatal error logs -XX:ErrorFile=/var/log/elasticsearch/hs_err_pid%p.log ## JDK 8 GC logging 8:-XX:+PrintGCDetails 8:-XX:+PrintGCDateStamps 8:-XX:+PrintTenuringDistribution 8:-XX:+PrintGCApplicationStoppedTime 8:-Xloggc:/var/log/elasticsearch/gc.log 8:-XX:+UseGCLogFileRotation 8:-XX:NumberOfGCLogFiles=32 8:-XX:GCLogFileSize=64m # JDK 9+ GC logging 9-:-Xlog:gc*,gc+age=trace,safepoint:file=/var/log/elasticsearch/gc.log:utctime,pid,tags:filecount=32,filesize=64m # due to internationalization enhancements in JDK 9 Elasticsearch need to set the provider to COMPAT otherwise # time/date parsing will break in an incompatible way for some date patterns and locals 9-:-Djava.locale.providers=COMPAT # temporary workaround for C2 bug with JDK 10 on hardware with AVX-512 10-:-XX:UseAVX=2
elasticsearch.yml文件
# ======================== Elasticsearch Configuration ========================= # # NOTE: Elasticsearch comes with reasonable defaults for most settings. # Before you set out to tweak and tune the configuration, make sure you # understand what are you trying to accomplish and the consequences. # # The primary way of configuring a node is via this file. This template lists # the most important settings you may want to configure for a production cluster. # # Please consult the documentation for further information on configuration options: # https://www.elastic.co/guide/en/elasticsearch/reference/index.html # # ---------------------------------- Cluster ----------------------------------- # # Use a descriptive name for your cluster: # #cluster.name: my-application # # ------------------------------------ Node ------------------------------------ # # Use a descriptive name for the node: # #node.name: node-1 # # Add custom attributes to the node: # #node.attr.rack: r1 # # ----------------------------------- Paths ------------------------------------ # # Path to directory where to store the data (separate multiple locations by comma): # path.data: /var/lib/elasticsearch # # Path to log files: # path.logs: /var/log/elasticsearch # # ----------------------------------- Memory ----------------------------------- # # Lock the memory on startup: # #bootstrap.memory_lock: true # # Make sure that the heap size is set to about half the memory available # on the system and that the owner of the process is allowed to use this # limit. # # Elasticsearch performs poorly when the system is swapping the memory. # # ---------------------------------- Network ----------------------------------- # # Set the bind address to a specific IP (IPv4 or IPv6): # #network.host: 192.168.0.1 # # Set a custom port for HTTP: # #http.port: 9200 # # For more information, consult the network module documentation. # # --------------------------------- Discovery ---------------------------------- # # Pass an initial list of hosts to perform discovery when new node is started: # The default list of hosts is ["127.0.0.1", "[::1]"] # #discovery.zen.ping.unicast.hosts: ["host1", "host2"] # # Prevent the "split brain" by configuring the majority of nodes (total number of master-eligible nodes / 2 + 1): # #discovery.zen.minimum_master_nodes: # # For more information, consult the zen discovery module documentation. # # ---------------------------------- Gateway ----------------------------------- # # Block initial recovery after a full cluster restart until N nodes are started: # #gateway.recover_after_nodes: 3 # # For more information, consult the gateway module documentation. # # ---------------------------------- Various ----------------------------------- # # Require explicit names when deleting indices: # #action.destructive_requires_name: true xpack.security.enabled: true
初步排查信息
初步调研认为内存、CPU飙升可能和未配置以下参数有关:
gateway.expected_nodes: 10 gateway.recover_after_time: 5m
集群状态查询结果
curl -XGET --user username:password http://localhost:9200/ { "name" : "web1.example.com", "cluster_name" : "elasticsearch", "cluster_uuid" : "S8fFQ993QDWkLY8lZtp_mQ", "version" : { "number" : "7.13.2", "build_flavor" : "default", "build_type" : "rpm", "build_hash" : "4d960a0733be83dd2543ca018aa4ddc42e956800", "build_date" : "2021-06-10T21:01:55.251515791Z", "build_snapshot" : false, "lucene_version" : "8.8.2", "minimum_wire_compatibility_version" : "6.8.0", "minimum_index_compatibility_version" : "6.0.0-beta1" }, "tagline" : "You Know, for Search" } curl --user username:password -sS http://localhost:9200/_cluster/health?pretty { "cluster_name" : "elasticsearch", "status" : "yellow", "timed_out" : false, "number_of_nodes" : 1, "number_of_data_nodes" : 1, "active_primary_shards" : 5, "active_shards" : 5, "relocating_shards" : 0, "initializing_shards" : 0, "unassigned_shards" : 4, "delayed_unassigned_shards" : 0, "number_of_pending_tasks" : 0, "number_of_in_flight_fetch" : 0, "task_max_waiting_in_queue_millis" : 0, "active_shards_percent_as_number" : 55.55555555555556 } curl --user username:password -sS http://localhost:9200/_cluster/allocation/explain?pretty { "index" : "example-amasty_product_1_v156", "shard" : 0, "primary" : false, "current_state" : "unassigned", "unassigned_info" : { "reason" : "INDEX_CREATED", "at" : "2021-09-14T16:52:28.854Z", "last_allocation_status" : "no_attempt" }, "can_allocate" : "no", "allocate_explanation" : "cannot allocate because allocation is not permitted to any of the nodes", "node_allocation_decisions" : [ { "node_id" : "2THEUTSaQdmOJAAhTTN71g", "node_name" : "web1.example.com", "transport_address" : "127.0.0.1:9300", "node_attributes" : { "ml.machine_memory" : "134622244864", "xpack.installed" : "true", "transform.node" : "true", "ml.max_open_jobs" : "512", "ml.max_jvm_size" : "51539607552" }, "node_decision" : "no", "weight_ranking" : 1, "deciders" : [ { "decider" : "same_shard", "decision" : "NO", "explanation" : "a copy of this shard is already allocated to this node" } ] } ] }
当前疑问与怀疑点
不清楚如何在单台机器上部署多节点,已知当前仅运行单节点,查阅资料显示需要3个主节点才能让集群状态变为green,想咨询单机器部署多节点的方法,以及是否需要增加数据节点。
主要怀疑点如下:
- 主/数据节点数量不足
- 垃圾回收器运行异常(已确认当前使用G1GC)
- 未配置崩溃恢复相关的
gateway.expected_nodes、gateway.recover_after_time参数
更新1:已收集到elasticsearch.log错误日志与_cluster/stats?pretty&human查询结果用于问题定位。
更新2:已找到限制副本数量的方法,可通过模板设置:
PUT _template/all { "template": "*", "settings": { "number_of_replicas": 0 } }
计划测试该配置是否能让集群状态变为green,暂不确定是否会对性能产生影响。其他正在调整的配置如下:
- 将JVM堆内存限制为31GB
- 文件描述符已设置为65535
- 最大线程数已设置为4096
- 虚拟内存上限已调整完成
max_map_count已提升至262144- 默认禁用G1GC
尝试将8-13:-XX:CMSInitiatingOccupancyFraction=75调整为8-13:-XX:CMSInitiatingOccupancyFraction=70,期望加快垃圾回收速度避免内存溢出,后续会上下调整该参数验证效果。同时计划尝试切换到G1GC,有相关案例显示该调整可解决类似的内存溢出问题。
更新3:完成上述调整后集群状态已变为green,夜间运行无异常,虽然性能不如配置50GB堆内存时的表现,但运行稳定。给后续Elasticsearch问题排查者的建议:优先完成官方引导检查项,可先达到基础性能要求。
更新4:后续发现JVM配置存在多路径冲突问题,系统管理员在/etc/elasticsearch/jvm.options.d目录下放置了heap_size.options配置了31GB堆内存,而主jvm.options文件配置为8GB,导致GC线程按8GB内存运行但实际占用31GB内存,删除冗余配置统一设置为31GB后情况有所稳定,但GC回收频率仍较高,新增索引属性时仍会出现GC内存溢出,只能删除索引重建解决,目前考虑重新部署Elasticsearch。
内容的提问来源于stack exchange,提问作者Kalvin Klien
相关产品推荐
相关产品推荐

