Vert.x Hazelcast集群中获取下线节点存储键及备份节点调用咨询
Hey there! Let's walk through how to address your scenario where you need to handle node failures, recover keys from the departed node, trigger actions on backup nodes, and run scheduled services. I'll break this down into actionable steps with code snippets tailored to your setup.
1. Correctly Retrieve Keys from the Departed Node
First off—you can't directly fetch keys from an offline node (it's unreachable once it's down). Instead, leverage Hazelcast's partition metadata to identify which partitions the departed node was responsible for, then pull those keys from the cluster's backups.
Here's how to do this in your MembershipListener:
@Override public void memberRemoved(MembershipEvent event) { Member departedNode = event.getMember(); PartitionService partitionService = hazelcastInstance.getPartitionService(); // Get all partitions that were assigned to the departed node Collection<Partition> nodePartitions = partitionService.getPartitions(departedNode); // Collect keys across all your IMap instances (adjust map names as needed) Set<Object> allDepartedKeys = new HashSet<>(); for (Partition partition : nodePartitions) { IMap<Object, Object> yourMap = hazelcastInstance.getMap("your-target-map"); // Use PartitionPredicate to fetch keys specific to this partition Set<Object> partitionKeys = yourMap.keySet(new PartitionPredicate(partition.getPartitionId())); allDepartedKeys.addAll(partitionKeys); } // Pass these keys to your recovery workflow initiateRecoveryWorkflow(allDepartedKeys); }
2. Trigger Actions on Backup Nodes
Hazelcast automatically promotes backup partitions to primary when a node goes down—so you don't need to manually "restore" data. But if you need to run custom logic on the node that now holds the data (the former backup), use these steps:
- For each key, find the current primary node using the partition service
- Use Vert.x's Event Bus or RPC to call services on that node
Example code to locate the backup/primary node and trigger actions:
private void invokeBackupNodeServices(Set<Object> keys) { PartitionService partitionService = hazelcastInstance.getPartitionService(); for (Object key : keys) { Partition partition = partitionService.getPartition(key); // Get the current primary node (this was the backup before the failure) Member primaryNode = partition.getOwner(); // Use Vert.x Event Bus to send a message to this node's service String nodeAddress = primaryNode.getAddress().getHost() + ":" + primaryNode.getAddress().getPort(); vertx.eventBus().send("backup-node-service." + nodeAddress, JsonObject.mapFrom(key), reply -> { if (reply.succeeded()) { // Handle successful response } else { // Handle failure } }); } }
3. Schedule Services at the Right Time
Timing is crucial here—Hazelcast's partition promotion happens asynchronously, so don't trigger your scheduler immediately after detecting a node failure. Wait for the cluster to stabilize first:
- Use a Vert.x timer to delay execution (adjust the delay based on your cluster size)
- Alternatively, register a
PartitionLostListenerto trigger logic once partitions are fully promoted
Example with a Vert.x timer:
private void initiateRecoveryWorkflow(Set<Object> keys) { // Wait 5 seconds for Hazelcast to complete partition promotion (tweak this value) vertx.setTimer(5000, timerId -> { // First, trigger backup node actions invokeBackupNodeServices(keys); // Then start your scheduled service startScheduledRecoveryService(keys); }); } private void startScheduledRecoveryService(Set<Object> keys) { // Example: Run a task every 10 minutes to verify key integrity vertx.setPeriodic(600000, periodicId -> { // Your scheduled logic here (e.g., validate keys, clean up stale data) validateRecoveredKeys(keys); }); }
Key Notes to Avoid Pitfalls
- Don't rely on direct node connections: Once a node is offline, you can't pull data from it—always use the cluster's partition metadata and backups.
- Adjust backup count: Ensure your Hazelcast config has a sufficient
backup-count(default is 1) to guarantee data availability during failover. - Test failover scenarios: Simulate node failures in staging to validate that your listener, key recovery, and scheduler logic work as expected.
内容的提问来源于stack exchange,提问作者Yogi

