Cassandra 3.9集群执行nodetool resetlocalschema后遇system_auth异常及GossipStage问题求助
system_auth Missing Error Hey there, let's work through this Cassandra issue step by step. I've tackled similar gossip backlog and schema sync problems with Cassandra 3.x clusters before, so here's a structured approach to get your node back to normal:
1. Fix the system_auth Keyspace Missing Error
Since you’ve enabled PasswordAuthenticator in cassandra.yaml, Cassandra relies on the system_auth keyspace to store user credentials. Running nodetool resetlocalschema likely wiped the local copy of this system keyspace’s metadata, triggering the missing error. Here’s how to restore it:
- First, log into a healthy node in your cluster and run this in
cqlshto get the full schema forsystem_auth:
Copy the entire output—this includes theDESCRIBE KEYSPACE system_auth;CREATE KEYSPACEstatement and all associatedCREATE TABLEcommands for user/role data. - If authentication blocks you from accessing the problematic node’s
cqlsh, temporarily switch theauthenticatorsetting in itscassandra.yamlfromPasswordAuthenticatortoAllowAllAuthenticator, then restart the node. - Once you can access
cqlshon the problematic node, paste and execute the copiedsystem_authschema statements to recreate the keyspace and its tables. - Switch the
authenticatorback toPasswordAuthenticatorincassandra.yaml, then restart the node again.
2. Resolve the GossipStage Backlog
The gossip backlog is often tied to schema sync failures or network connectivity issues between the node and the rest of the cluster. Now that system_auth is fixed, let’s address the gossip problem:
- First, verify network connectivity: Ensure firewall rules allow traffic on Cassandra’s gossip port (7000) and CQL port (9042) between the problematic node and all other cluster nodes.
- Check the node’s status from a healthy node with:
If the node is markednodetool statusDN(Down) orUJ(Up/Joining), we’ll need to rejoin it properly:- Stop the problematic node:
sudo service cassandra stop - Clear the node’s local data (safe for multi-node clusters, as data will sync back via replication):
sudo rm -rf /var/lib/cassandra/data/* sudo rm -rf /var/lib/cassandra/commitlog/* sudo rm -rf /var/lib/cassandra/saved_caches/* - Double-check
cassandra.yamlon the node to ensureseed_providerpoints to 2-3 healthy seed nodes in your cluster. - Start the node again:
sudo service cassandra start - Monitor the system logs (
/var/log/cassandra/system.log) to confirm the node starts joining the cluster and gossip messages process without backlog.
- Stop the problematic node:
- After the node is up, run this to verify gossip is working:
You should see entries for all other cluster nodes in the output.nodetool gossipinfo
3. Verify Node Health
Once the node is back online, confirm it’s fully recovered:
- Run
nodetool statusto check if the node is markedUN(Up/Normal). - Use
nodetool tpstatsto check theGossipStage—there should be no pending tasks or backlogged messages. - In
cqlsh, runDESCRIBE KEYSPACE system_authto confirm the keyspace exists and has the correct tables. - Test authentication by logging into
cqlshwith a valid user to ensurePasswordAuthenticatorworks as expected.
内容的提问来源于stack exchange,提问作者Varsha

