Hazelcast中央服务器数据变更时应用Near Cache未失效问题
Hazelcast Near Cache 事件丢失问题
我们在Kubernetes 1.21集群中部署了Hazelcast 3.12.4及其他应用。正常情况下,当任一应用实例更新中央Hazelcast缓存时,会触发事件并发送给所有存储该数据Near Cache的应用实例;应用收到事件后会失效Near Cache数据,从而从中央缓存获取最新数据。但目前随机出现部分应用实例未接收或未处理Hazelcast事件的情况,导致这些实例的Near Cache数据未更新,出现部分实例缓存最新数据、部分实例缓存过期数据的异常,且无固定复现步骤。
客户端配置文件
<hazelcast-client xmlns="http://www.hazelcast.com/schema/client-config" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.hazelcast.com/schema/client-config http://www.hazelcast.com/schema/client-config/hazelcast-client-config-5.1.xsd"> <cluster-name>dev</cluster-name> <instance-name>UPC</instance-name> <properties> <property name="hazelcast.client.shuffle.member.list">true</property> <property name="hazelcast.client.heartbeat.timeout">600000</property> <property name="hazelcast.client.heartbeat.interval">180000</property> <property name="hazelcast.client.event.queue.capacity">1000000</property> <property name="hazelcast.client.invocation.timeout.seconds">120</property> <property name="hazelcast.client.statistics.enabled">false</property> <property name="hazelcast.discovery.enabled">false</property> <property name="hazelcast.invalidation.reconciliation.interval.seconds">31</property> <property name="hazelcast.invalidation.max.tolerated.miss.count">1</property> <property name="hazelcast.map.invalidation.batch.enabled">true</property> <property name="hazelcast.map.invalidation.batch.size">5</property> <property name="hazelcast.map.invalidation.batchfrequency.seconds">5</property> </properties> <network> <cluster-members> <address>10.151.32.200:31050</address> </cluster-members> <discovery-strategies> <discovery-strategy enabled="false" class="com.hazelcast.kubernetes.HazelcastKubernetesDiscoveryStrategy"> <properties> <!-- configure discovery service API lookup --> <property name="service-name">hazelcast</property> <property name="namespace">default</property> </properties> </discovery-strategy> </discovery-strategies> <smart-routing>true</smart-routing> <redo-operation>true</redo-operation> <connection-timeout>90000</connection-timeout> </network> <near-cache name="default"> <in-memory-format>OBJECT</in-memory-format> <serialize-keys>true</serialize-keys> <invalidate-on-change>true</invalidate-on-change> <eviction eviction-policy="NONE" max-size-policy="ENTRY_COUNT"/> </near-cache> <connection-strategy async-start="false" reconnect-mode="ASYNC"> <connection-retry> <initial-backoff-millis>2000</initial-backoff-millis> <max-backoff-millis>60000</max-backoff-millis> <multiplier>3</multiplier> <cluster-connect-timeout-millis>5000</cluster-connect-timeout-millis> <jitter>0.5</jitter> </connection-retry> </connection-strategy> </hazelcast-client>
排查与优化建议
- 调优心跳参数:当前心跳间隔180秒、超时600秒过长,Kubernetes网络易出现短暂波动,建议将
hazelcast.client.heartbeat.interval改为30000(30秒),hazelcast.client.heartbeat.timeout改为120000(120秒),确保客户端及时感知集群连接状态,避免失联后事件接收中断。 - 监控事件队列状态:开启客户端统计(将
hazelcast.client.statistics.enabled设为true),监控client.event.queue.size指标,排查是否因事件处理速度跟不上导致队列溢出丢事件。若有积压,需优化客户端事件处理逻辑,或适当调整队列容量。 - 调整失效批处理配置:当前失效批大小5、频率5秒的设置,在更新频繁场景下易导致事件延迟或丢包。建议将
hazelcast.map.invalidation.batch.size调至50-100,hazelcast.map.invalidation.batchfrequency.seconds改为1秒,平衡事件传输的延迟与吞吐量。 - 启用Kubernetes集群发现:当前客户端仅绑定单个集群成员地址,若该成员故障或网络分区,客户端将无法接收事件。建议启用Kubernetes发现策略(将
discovery-strategy的enabled设为true,移除cluster-members固定地址配置),让客户端自动发现所有集群成员,提升连接可靠性。 - 加快失效 reconciliation 速度:当前
hazelcast.invalidation.reconciliation.interval.seconds为31秒,若事件丢失,过期数据修正速度慢。建议调至10秒,同时保持hazelcast.invalidation.max.tolerated.miss.count为1,确保错过少量事件后能快速触发数据同步。 - 排查网络通信问题:检查Kubernetes网络策略是否限制了客户端与Hazelcast集群的通信,通过集群监控工具或tcpdump排查节点间是否存在网络丢包,确认事件传输链路的稳定性。
内容的提问来源于stack exchange,提问作者Safvan Kothawala
相关产品推荐
相关产品推荐

