Hazelcast IMap.tryLock设100ms超时却阻塞超1分钟问题求助
Alright, let's dig into this Hazelcast lock timeout issue you're facing. First, let's recap your environment and setup to make sure we're aligned:
- Hazelcast Cluster: 3.9.2, 2 nodes running on Windows Server 2012 R2 Standard (Oracle JDK 1.8.0_144)
- Clients: 2 to 20 instances on separate VMs (3.10.0-327.28.3.el7.x86_64, IBM JDK 1.7.1_64)
Cluster-Side (hazelcast.xml)
<map name="lock*"> <in-memory-format>BINARY</in-memory-format> <statistics-enabled>true</statistics-enabled> <backup-count>1</backup-count> <eviction-policy>NONE</eviction-policy> </map>
Client-Side (hazelcast-client.xml)
<hazelcast-client xmlns="http://www.hazelcast.com/schema/client-config" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.hazelcast.com/schema/client-config file:///C:/caching/hazelcast-client-config-3.9.xsd"> <group> <name>OUR_GROUP_NAME</name> </group> <properties> <property name="hazelcast.client.shuffle.member.list">true</property> <property name="hazelcast.client.heartbeat.timeout">60000</property> <property name="hazelcast.client.heartbeat.interval">5000</property> <property name="hazelcast.client.event.thread.count">10</property> <property name="hazelcast.client.event.queue.capacity">1000000</property> <property name="hazelcast.client.invocation.timeout.seconds">35</property> <property name="hazelcast.client.statistics.enabled">true</property> </properties> <network> <cluster-members> <address>tvlcacheqa1.blqa.qa:5709</address> <address>tvlcacheqa2.blqa.qa:5709</address> </cluster-members> <smart-routing>true</smart-routing> <redo-operation>true</redo-operation> <connection-attempt-period>15000</connection-attempt-period> <connection-attempt-limit>1048576</connection-attempt-limit> <socket-options> <tcp-no-delay>false</tcp-no-delay> <keep-alive>true</keep-alive> <reuse-address>true</reuse-address> <linger-seconds>5</linger-seconds> <timeout>-1</timeout> <buffer-size>64</buffer-size> </socket-options> </network> <near-cache name="cache*"> <in-memory-format>OBJECT</in-memory-format> <invalidate-on-change>true</invalidate-on-change> <time-to-live-seconds>1800</time-to-live-seconds> <max-idle-seconds>1800</max-idle-seconds> <eviction eviction-policy="LRU" max-size-policy="ENTRY_COUNT" size="10000"/> </near-cache> </hazelcast-client>
Note: All unmentioned configurations use default values, and no programmatic config overrides are applied.
Batch Processing Concurrency Pattern
Clients execute this logic in parallel:
synchronizer.startSyncSection(key, 100); try { doSomeCriticalStuff(); } finally { synchronizer.endSyncSection(key); }
Synchronizer Core (IMap-Based Locking)
@Override public void startSynchedSection(MultiKey<?> key, long tryLockTimeoutInMs, long releaseLockTimeoutInMs) { keyNullCheck(key); tryLockTimeoutInMs = Math.max(tryLockTimeoutInMs, minimumObtainLockTimeoutInMs); if (isClusterReady()) { boolean locked = false; try { locked = this.locks.tryLock(key, tryLockTimeoutInMs, TimeUnit.MILLISECONDS, releaseLockTimeoutInMs, TimeUnit.MILLISECONDS); } catch (InterruptedException e) { throw new TechnicalException(e); } if (!locked) { throw new SyncTimeoutException(FAILED_TO_OBTAIN_EXCEPTION + key); } int lockCounter = incrementLockCounter(key); } else { throw new SyncTimeoutException(CLUSTER_NOT_READY_EXCEPTION); } } /** Note This should be called in a FINALLY section!!! */ @Override public void endSynchedSection(MultiKey<?> key) { keyNullCheck(key); int lockCounterBefore = getThreadLocalCounter(key.toString()).get(); if (lockCounterBefore == 0) { return; } try { int lockCounterAfter = decrementLockCounter(key); if (this.locks.isLocked(key)) { this.locks.unlock(key); } } catch (OperationTimeoutException e) { this.logger.warn("endSynchedSection - Lock-> {} was not released properly in Hazelcast because of exception:\n{}\n in Thread={}", key, e .getMessage(), Thread.currentThread().getName()); } }
Where this.locks is an IMap<String, String> instance.
During batch runs, threads frequently block on the IMap.tryLock call. Even though you set tryLockTimeoutInMs=100ms, threads can block for up to 2 minutes. Complicating things, you can't reproduce this behavior in your test environment.
Let's walk through the most likely culprits and actionable steps to diagnose this:
1. Mismatched JDK Versions & Serialization
Your cluster uses Oracle JDK 1.8 while clients use IBM JDK 1.7. While Hazelcast 3.9.2 supports both, IBM JDK has subtle differences in threading and serialization:
- Verify that your
MultiKeyclass serializes consistently across both JDKs. Mismatched serialization could lead to unexpected lock key collisions, where multiple threads think they're locking different keys but actually target the same one. - Check IBM JDK's concurrency library behavior—some implementations handle thread waits differently under high load, which could affect lock timeout adherence.
2. Network Latency & TCP Configuration
Your client socket config has <tcp-no-delay>false</tcp-no-delay>, which enables Nagle's algorithm. This delays small packet transmission (like lock requests) to optimize throughput, but under high load, this can cause significant latency that makes your 100ms timeout irrelevant.
- Quick fix: Set
<tcp-no-delay>true</tcp-no-delay>to disable Nagle's algorithm and reduce latency for small, time-sensitive requests. - Monitor network metrics between clients and cluster nodes: look for packet loss, high round-trip times (RTT), or jitter. Even occasional spikes in RTT can cause lock requests to exceed the timeout.
3. Cluster Node Load & GC Pauses
If cluster nodes are under heavy CPU, memory, or I/O load, they may not process lock requests in a timely manner:
- Enable detailed GC logging on cluster nodes. Long stop-the-world (STW) GC pauses can make the cluster unresponsive temporarily, leading to lock requests taking far longer than the configured timeout.
- Use Hazelcast Management Center to monitor lock statistics (acquisition rates, wait times, pending requests) and node resource usage (CPU, memory).
4. Invocation Timeout vs Lock Timeout Misalignment
Your client's hazelcast.client.invocation.timeout.seconds is set to 35 seconds—this is a critical setting! The invocation timeout wraps the entire lock request process. If the lock request gets stuck in the client's queue or the cluster doesn't respond, the client will wait up to 35 seconds before timing out, ignoring your 100ms lock-specific timeout.
- Hazelcast 3.9.x doesn't support per-invocation timeouts natively, but you can wrap the
tryLockcall in aCompletableFuturewith a 100ms timeout to enforce your desired wait time.
5. Thread-Local Lock Counter Bugs
Your synchronizer uses a thread-local counter to track lock acquisitions, but there's a potential flaw in the endSynchedSection logic:
- If a thread acquires the lock multiple times (counter increments multiple times), calling
unlockonce will release the lock immediately—even if the thread still needs it. This can lead to race conditions and orphaned locks. - Conversely, if the counter isn't decremented correctly,
endSynchedSectionmight skip callingunlockentirely, leaving the lock held indefinitely and causing permanent contention for that key. - Audit the
incrementLockCounteranddecrementLockCountermethods to ensure they correctly track acquisitions/releases, and thatunlockis called exactly once per successful lock acquisition.
6. Test vs Production Environment Differences
Since you can't reproduce this in test, look for key gaps:
- Load: Production likely has higher concurrency, more unique keys, or longer-running
doSomeCriticalStuff()operations that increase lock contention. - Network: Test environments often have lower latency and no packet loss, which masks network-related issues.
- JVM Configs: Check if GC settings, heap sizes, or other JVM parameters differ between test and production (e.g., production might use a different GC collector with longer pauses).
7. Hazelcast 3.9.2 Known Issues
Hazelcast 3.9.2 is an older version (released in 2018) with known bugs related to lock handling:
- Some versions had issues with lock requests getting stuck during cluster rebalancing or node load spikes.
- Consider upgrading to the latest patch version of 3.9.x (e.g., 3.9.6) or a newer major version, as many lock-related bugs have been fixed since 3.9.2.
- Fix TCP No-Delay: Update the client config to set
<tcp-no-delay>true</tcp-no-delay>to reduce network latency for lock requests. - Enforce Lock Timeout: Wrap the
tryLockcall in aCompletableFuturewith a 100ms timeout to bypass the global invocation timeout. - Audit Lock Counter Logic: Verify that every lock acquisition is paired with exactly one release, and that the thread-local counter works as intended.
- Monitor Cluster Metrics: Enable GC logging on cluster nodes and use Hazelcast Management Center to track lock statistics and node health.
- Simulate Production Load: Replicate production-level concurrency, key volume, and network conditions in your test environment to reproduce the issue.
内容的提问来源于stack exchange,提问作者Maxim Duhovniy

