Hazelcast cluster.timeDiff指标含义及相关集群时钟同步问题咨询
cluster.timeDiff Metric (Version 3.8+) First, let's recap the relevant HealthMonitor output you shared:
.HealthMonitor; [xxx/com/hazelcast/internal/diagnostics/HealthMonitor]; [3.8] processors=12, physical.memory.total=31.1G, physical.memory.free=1.2G, swap.space.total=15.7G, swap.space.free=15.7G, heap.memory.used=181.1M, heap.memory.free=74.9M, heap.memory.total=256.0M, heap.memory.max=256.0M, heap.memory.used/total=70.75%, heap.memory.used/max=70.75% , minor.gc.count=2626, minor.gc.time=27401ms, major.gc.count=0, major.gc.time=0ms, load.process=0.00%, load.system=0.38%, load.systemAverage=549.00%, thread.count=91, thread.peakCount=177, cluster.timeDiff=-242516 , event.q.size=0, executor.q.async.size=0, executor.q.client.size=0, executor.q.query.size=0, executor.q.scheduled.size=0, executor.q.io.size=0, executor.q.system.size=0, executor.q.operations.size=0, executor.q.priorityOperation.size=0, operations.completed.count=15118896, executor.q.mapLoad.size=0, executor.q.mapLoadAllKeys.size=0, executor.q.cluster.size=0, executor.q.response.size=0, operations.running.count=0, operations.pending.invocations.percentage=0.00%, operations.pending.invocations.count=0, proxy.count=0, clientEndpoint.count=0, connection.active.count=20, client.connection.count=0, connection.count=15;
And the critical log entry:
2021-03-06T08:16:39,495; WARN; [xxx/com/hazelcast/internal/cluster/impl/ClusterHeartbeatManager]; [x.x.x.x]:5701 [xxx] [3.8] Ignoring heartbeat from Member [x.x.x.x]:5724 - 0a519310-ac9f-4085-b38d-aa7a7ec8ff1f lite since it is expired (now: 2021-03-06 08:12:36.979, timestamp: 2021-03-06 08:11:32.385);
Now let's answer your questions directly:
1. What does the cluster.timeDiff metric represent?
cluster.timeDiff measures the time difference (in milliseconds) between the current Hazelcast member's system clock and another member's clock in the cluster. The negative value (-242516ms, ~-4 minutes) in your output means the current member's clock is 4 minutes behind the member it's comparing against. Your log confirms this: the JVM's log timestamp (08:16:39) is 4 minutes later than the "now" timestamp reported by the member (08:12:36), matching the absolute value of cluster.timeDiff.
2. Does this metric indicate missing NTP synchronization?
Absolutely. A time difference of this magnitude (~4 minutes) is a clear sign that your cluster members are not synchronized via NTP (or any clock synchronization protocol). Hazelcast relies on consistent system clocks to manage critical cluster operations like heartbeats, session timeouts, and member lifecycle. The issues you're seeing—ignored expired heartbeats, clock jump warnings, and members being removed from the cluster—are direct consequences of this clock inconsistency.
3. Is cluster.timeDiff the max difference between any two members, or between the current member and others?
Based on Hazelcast 3.8's internal logic (evidenced by the ClusterHeartbeatManager logs), cluster.timeDiff represents the maximum time difference between the current member and all other members in the cluster. It's not a global maximum across any pair of members; it's calculated from the perspective of the member generating the HealthMonitor output, taking the largest delta between its clock and every other member's clock.
内容的提问来源于stack exchange,提问作者Haphil

