执行nodetool decommission后Cassandra集群节点状态异常求助
Hey there, let's break down why you're seeing mixed UD/UN statuses for your remaining Cassandra nodes after decommissioning, and what steps you might have missed (plus how to fix the current state).
First, let's clarify the status codes
Just to make sure we're on the same page:
UN: Up/Normal (the node is online and functioning as expected)UD: If you’re seeing this for nodes that weren’t decommissioned, it likely means a gossip sync issue where some nodes incorrectly mark those nodes as decommissioned. Alternatively, it might be a typo forDN(Down/Normal)—where some nodes see the target as offline while others see it as online. Either way, root causes often boil down to gossip or network issues.
Why this happens (and how to fix it)
1. Gossip Protocol Sync Delays or Inconsistencies
Cassandra relies entirely on gossip to share node status across the cluster. After decommissioning a node, it can take time for this update to propagate—especially in larger clusters or if there’s minor network latency. Sometimes, nodes get stuck with stale gossip data.
Fixes:
- Wait 10-15 minutes for gossip to fully sync, then re-run
nodetool statuson all nodes. - On nodes showing incorrect statuses, run
nodetool gossipinfoand look for the problematic node’s entry. Compare it against a node showing the correctUNstatus to spot discrepancies in theSTATUSfield. - Force a gossip reset on the misbehaving nodes: run
nodetool disablegossip, wait 30 seconds, thennodetool enablegossipto trigger a fresh sync.
2. Partial Network Partition
If your cluster has a network split (even a temporary one), nodes on one side won’t communicate with the other. This leads to inconsistent status views—nodes in one partition see others as down/decommissioned, while the other partition sees them as normal.
Fixes:
- Verify network connectivity between conflicting nodes: ping the target node from the one showing
UD, and check if Cassandra’s gossip port (7000) and CQL port (9042) are reachable withtelnet <node-ip> 7000. - Check the
system.logon misbehaving nodes for errors likeCannot gossip with ...orConnection refused—these confirm network issues. - Resolve any firewall rules, routing issues, or network outages blocking communication.
3. Incomplete Decommission Process
Even if nodetool decommission seemed to finish successfully, there might have been hidden failures during data streaming that left the cluster inconsistent. For example, a partial stream failure could make some nodes think the decommissioned node is still active.
Fixes:
- Check the decommissioned node’s
system.logfor errors likeStream failedorRange move pendingto confirm the process completed cleanly. - On remaining nodes, run
nodetool netstatsto ensure no pending stream requests or range moves exist. - Run
nodetool cleanupon all remaining nodes to clear token ranges that should no longer belong to them (this also frees up disk space).
Decommission Best Practices to Avoid This Issue
To prevent status inconsistencies in future decommissions, follow these steps:
Before Decommissioning
- Confirm the node to remove is in
UNstate withnodetool status—never decommission a node that’s down or in an error state. - Check cluster health: run
nodetool tpstatsandnodetool netstatsto ensure no pending tasks or stream failures. - Ensure your replication factor (RF) matches the remaining cluster size. For example, if RF=3, you need at least 3 nodes left to maintain quorum—otherwise, decommission will fail or leave the cluster unstable.
- Disable auto-restart for the node (e.g., via systemd) to prevent accidental rejoin after shutdown.
During Decommissioning
- Don’t interrupt the
nodetool decommissioncommand—wait until it explicitly outputsDecommissioned successfully. - Monitor data streaming with
nodetool netstatsuntilPending range movesdrops to 0 and all stream requests are marked completed. - Keep an eye on the node’s
system.logfor streaming errors.
After Decommissioning
- Stop the decommissioned node’s Cassandra service with
nodetool stopdaemonto fully take it offline. - Run
nodetool statuson all remaining nodes to confirm consistent statuses (all should beUN, and the decommissioned node should showDD). - If inconsistencies appear immediately, wait for gossip sync first before troubleshooting further.
If you’ve tried all the above and still see mixed statuses, a last resort is to restart the nodes showing incorrect statuses—just be sure to drain them first with nodetool drain to avoid data loss.
内容的提问来源于stack exchange,提问作者Avis

