You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Hazelcast 5.2.3三节点集群分区迁移异常及无分区节点排查

问题描述

在Windows Server 2016系统上搭建3节点Hazelcast 5.2.3社区版集群,集群已形成但有一个节点未分配任何分区,客户端日志抛出OperationTimeoutException,根源为PartitionMigratingException。客户端版本同样为5.2.3,集群配置如下:

hazelcast:
  listeners:
    - com.keshettv.keshetcoreinfra.service.cache.hazelcast.manager.impl.ClusterMembershipListener
  cluster-name: development
  properties:
    hazelcast.jmx: true
    hazelcast.socket.connect.timeout.seconds: 10
    hazelcast.logging.type: log4j
    hazelcast.jet.enabled: true
  network:
    reuse-address: true
    port:
      auto-increment: true
      port: 5701
    outbound-ports:
      ports: 34500
    join:
      auto-detection:
        enabled: false
      tcp-ip:
        enabled: true
        member-list:
          - [replaced the ip list]
        connection-timeout-seconds: 300
    interfaces:
      enabled: true
      interfaces: 
        - [replaced the ip list]
    ssl:
      enabled: false
      properties:
        protocol: TLSv1.2
        mutualAuthentication: REQUIRED
        keyStore: /opt/hazelcast.keystore
        keyStorePassword: secret.97531
        keyStoreType: PKCS12
        trustStore: /opt/hazelcast.truststore
        trustStorePassword: changeit
        trustStoreType: PKCS12
        keyMaterialDuration: PT10M
    failure-detector:
      icmp:
        enabled: false
        timeout-milliseconds: 1000
        fail-fast-on-startup: true
        interval-milliseconds: 1000
        max-attempts: 2
        parallel-mode: true
        ttl: 255
    symmetric-encryption:
      enabled: false
      algorithm: PBEWithMD5AndDES
      salt: thesalt
      password: thepass
      iteration-count: 19
  executor-service:
    default:
      statistics-enabled: true
      pool-size: 16
      queue-capacity: 0
  durable-executor-service:
    default:
      pool-size: 16
      durability: 1
      capacity: 100
  scheduled-executor-service:
    default:
      pool-size: 16
      durability: 1
      capacity: 100
      capacity-policy: PER_NODE
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
  set:
    default:
      statistics-enabled: false
      backup-count: 1
      async-backup-count: 0
      max-size: 10
  queue:
    default:
      statistics-enabled: true
      max-size: 0
      backup-count: 1
      async-backup-count: 0
      empty-queue-ttl: -1
      queue-store:
        class-name: com.hazelcast.QueueStoreImpl
        properties:
          binary: false
          memory-limit: 1000
          bulk-load: 500
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
  map:
    default:
      in-memory-format: BINARY
      metadata-policy: CREATE_ON_UPDATE
      statistics-enabled: true
      per-entry-stats-enabled: false
      cache-deserialized-values: ALWAYS
      backup-count: 0
      async-backup-count: 0
      time-to-live-seconds: 0
      max-idle-seconds: 0
      eviction:
        eviction-policy: LRU
        max-size-policy: PER_NODE
        size: 0
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
      read-backup-data: false
      merkle-tree:
        enabled: false
        depth: 10
      event-journal:
        enabled: false
        capacity: 10000
        time-to-live-seconds: 0
    OBJECTS_CACHE:
      in-memory-format: BINARY
      metadata-policy: CREATE_ON_UPDATE
      statistics-enabled: true
      per-entry-stats-enabled: false
      cache-deserialized-values: NEVER
      backup-count: 1
      async-backup-count: 0
      time-to-live-seconds: 0
      max-idle-seconds: 0
      eviction:
        eviction-policy: LRU
        max-size-policy: USED_HEAP_PERCENTAGE
        size: 10
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
      read-backup-data: false
      near-cache:
        in-memory-format: OBJECT
        invalidate-on-change: false
        time-to-live-seconds: 60
        eviction:
          eviction-policy: LRU
          max-size-policy: ENTRY_COUNT
          size: 1000
        cache-local-entries: true
      map-store:
        enabled: true
        initial-mode: LAZY
        class-name: com.keshettv.keshetcoreinfra.service.cache.hazelcast.manager.impl.HtmlMapStore
        write-delay-seconds: 60
        write-batch-size: 10000
        write-coalescing: true
        properties:
          connection-string: [Some connection string]
          database-name: hazelcast
          collection-name: OBJECTS_CACHE
          connections-per-host: 50
          min-connections-per-host: 10
          max-connection-idle-time: 60000
          max-connection-life-time: 120000
          max-wait-time: 5000
    OBJECTS_CACHE_only_new:
      in-memory-format: BINARY
      metadata-policy: CREATE_ON_UPDATE
      statistics-enabled: true
      per-entry-stats-enabled: false
      cache-deserialized-values: NEVER
      backup-count: 1
      async-backup-count: 0
      time-to-live-seconds: 0
      max-idle-seconds: 0
      eviction:
        eviction-policy: LRU
        max-size-policy: USED_HEAP_PERCENTAGE
        size: 10
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
      read-backup-data: false
      map-store:
        enabled: true
        initial-mode: LAZY
        class-name: com.keshettv.keshetcoreinfra.service.cache.hazelcast.manager.impl.HtmlMapStore
        write-delay-seconds: 60
        write-batch-size: 10000
        write-coalescing: true
        properties:
          connection-string: [Some connection string]
          database-name: hazelcast
          collection-name: OBJECTS_CACHE_only_new
          connections-per-host: 50
          min-connections-per-host: 10
          max-connection-idle-time: 60000
          max-connection-life-time: 120000
          max-wait-time: 5000
    AXIS_CACHE:
      in-memory-format: BINARY
      metadata-policy: CREATE_ON_UPDATE
      statistics-enabled: true
      per-entry-stats-enabled: false
      cache-deserialized-values: NEVER
      backup-count: 1
      async-backup-count: 0
      time-to-live-seconds: 0
      max-idle-seconds: 0
      eviction:
        eviction-policy: LRU
        max-size-policy: PER_NODE
        size: 1000
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
      read-backup-data: false
      near-cache:
        in-memory-format: OBJECT
        invalidate-on-change: false
        time-to-live-seconds: 60
        eviction:
          eviction-policy: LRU
          max-size-policy: ENTRY_COUNT
          size: 1000
        cache-local-entries: true
      map-store:
        enabled: true
        initial-mode: LAZY
        class-name: com.keshettv.keshetcoreinfra.service.cache.hazelcast.manager.impl.HtmlMapStore
        write-delay-seconds: 60
        write-batch-size: 1000
        write-coalescing: true
        properties:
          connection-string: [Some connection string]
          database-name: hazelcast
          collection-name: AXIS_CACHE
          connections-per-host: 50
          min-connections-per-host: 10
          max-connection-idle-time: 60000
          max-connection-life-time: 120000
          max-wait-time: 5000
    AXIS_CACHE_only_new:
      in-memory-format: BINARY
      metadata-policy: CREATE_ON_UPDATE
      statistics-enabled: true
      per-entry-stats-enabled: false
      cache-deserialized-values: NEVER
      backup-count: 1
      async-backup-count: 0
      time-to-live-seconds: 0
      max-idle-seconds: 0
      eviction:
        eviction-policy: LRU
        max-size-policy: PER_NODE
        size: 1000
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
      read-backup-data: false
      map-store:
        enabled: true
        initial-mode: LAZY
        class-name: com.keshettv.keshetcoreinfra.service.cache.hazelcast.manager.impl.HtmlMapStore
        write-delay-seconds: 60
        write-batch-size: 1000
        write-coalescing: true
        properties:
          connection-string: [Some connection string]
          database-name: hazelcast
          collection-name: AXIS_CACHE_only_new
          connections-per-host: 50
          min-connections-per-host: 10
          max-connection-idle-time: 60000
          max-connection-life-time: 120000
          max-wait-time: 5000
    HTML_CACHE:
      in-memory-format: BINARY
      metadata-policy: CREATE_ON_UPDATE
      statistics-enabled: true
      per-entry-stats-enabled: false
      cache-deserialized-values: NEVER
      backup-count: 1
      async-backup-count: 0
      time-to-live-seconds: 0
      max-idle-seconds: 0
      eviction:
        eviction-policy: LRU
        max-size-policy: USED_HEAP_PERCENTAGE
        size: 10
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
      read-backup-data: false
      near-cache:
        in-memory-format: OBJECT
        invalidate-on-change: false
        time-to-live-seconds: 60
        eviction:
          eviction-policy: LRU
          max-size-policy: ENTRY_COUNT
          size: 1000
        cache-local-entries: true
      map-store:
        enabled: true
        initial-mode: LAZY
        class-name: com.keshettv.keshetcoreinfra.service.cache.hazelcast.manager.impl.HtmlMapStore
        write-delay-seconds: 60
        write-batch-size: 10000
        write-coalescing: true
        properties:
          connection-string: [Some connection string]
          database-name: hazelcast
          collection-name: HTML_CACHE
          connections-per-host: 50
          min-connections-per-host: 10
          max-connection-idle-time: 60000
          max-connection-life-time: 120000
          max-wait-time: 5000
    HTML_CACHE_only_new:
      in-memory-format: BINARY
      metadata-policy: CREATE_ON_UPDATE
      statistics-enabled: true
      per-entry-stats-enabled: false
      cache-deserialized-values: NEVER
      backup-count: 1
      async-backup-count: 0
      time-to-live-seconds: 0
      max-idle-seconds: 0
      eviction:
        eviction-policy: LRU
        max-size-policy: USED_HEAP_PERCENTAGE
        size: 10
      merge-policy:
        batch-size: 100
        class-name: PutIfAbsentMergePolicy
      read-backup-data: false
      map-store:
        enabled: true
        initial-mode: LAZY
        class-name: com.keshettv.keshetcoreinfra.service.cache.hazelcast.manager.impl.HtmlMapStore
        write-delay-seconds: 60
        write-batch-size: 10000
        write-coalescing: true
        properties:
          connection-string: [Some connection string]
          database-name: hazelcast
          collection-name: HTML_CACHE_only_new
          connections-per-host: 50
          min-connections-per-host: 10
          max-connection-idle-time: 60000
          max-connection-life-time: 120000
          max-wait-time: 5000
原因分析
  • 分区迁移卡住:节点加入集群时,分区迁移流程因资源不足、网络延迟或存储层阻塞未完成,导致该节点一直处于等待接收分区的状态。客户端请求因等待分区迁移完成超时,触发PartitionMigratingException,最终引发OperationTimeoutException。
  • 资源瓶颈:目标节点的CPU、内存或磁盘IO不足,无法处理分区迁移的大量数据传输与存储操作,拖慢迁移进度。
  • MapStore影响:多个Map配置了自定义HtmlMapStore且write-delay-seconds为60,若后端存储(如数据库)响应缓慢或连接池资源耗尽,会占用节点大量资源,间接阻碍分区迁移。
  • 网络配置限制:outbound-ports仅指定单个端口34500,可能因端口资源不足导致节点间分区迁移的数据包传输受阻;节点间网络存在隐性延迟,超出默认的分区迁移超时阈值。
  • 集群状态不稳定:节点加入时集群未完全稳定,导致Hazelcast的分区分配逻辑未正常触发或执行失败,造成节点无分区分配。
解决办法

紧急恢复操作

  • 重启无分区节点:优雅关闭该节点,待集群重新平衡后再启动,触发分区重新分配逻辑。
  • 强制分区平衡:若部署了Hazelcast Management Center,直接执行分区重新平衡操作;或在代码中调用hazelcastInstance.getPartitionService().forceLocalPartitionBalance()方法触发平衡。

配置优化

  1. 调整分区迁移参数
    在配置的properties节点添加以下参数,延长迁移超时时间并增加节点加入后的迁移延迟:
properties:
  # 其他原有参数...
  hazelcast.partition.migration.timeout.seconds: 300
  hazelcast.partition.migration.delay.millis: 5000
  1. 优化网络配置
  • 取消outbound-ports的单端口限制,改为端口范围,避免端口不足影响节点间通信:
outbound-ports:
  ports: 34500-34600
  • 检查Windows Server防火墙规则,确保节点间的5701+端口及迁移所需端口完全开放。
  1. 资源与存储层优化
  • 检查无分区节点的CPU、内存使用率,确保资源充足;若内存不足,调整JVM堆大小。
  • 优化HtmlMapStore的数据库连接池配置:降低connections-per-host数值(如从50调整为20),避免占用过多节点资源;同时检查后端数据库性能,确保存储层响应及时。
  1. 调整Map配置
  • 对于启用MapStore的Map,暂时将initial-mode改为EAGER,确保节点启动时优先完成数据加载,避免迁移时的资源冲突;或临时禁用MapStore,待分区稳定后再启用。
  • 将near-cache的invalidate-on-change设为true,避免缓存不一致导致的额外资源消耗。
  1. 节点健康检查
  • 查看节点日志,排查是否存在GC频繁、线程阻塞等JVM层面的异常;
  • 启用Hazelcast健康检查机制,确认节点状态正常。

内容的提问来源于stack exchange,提问作者Ohad Behore

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 19:02:00