You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Cassandra 3.0.5集群节点启动失败及多异常问题求助

Cassandra 3.0.5双DC集群启动及Schema一致性故障排查与恢复

我们有一个Cassandra 3.0.5版本的双DC集群(共约30节点,18个节点在一个DC,其余在另一个DC),近期因GC停顿问题调整了所有节点的JVM参数(修改MAX_HEAP_SIZE),准备滚动重启生效。第一个节点重启正常,但第二个节点关机后无法启动,后续出现一系列连锁问题。

一、节点启动失败错误

第二个节点重启时抛出以下异常:

INFO  07:45:34 Initializing system_schema.keyspaces
INFO  07:45:34 Initializing system_schema.tables
INFO  07:45:34 Initializing system_schema.columns
INFO  07:45:34 Initializing system_schema.triggers
INFO  07:45:34 Initializing system_schema.dropped_columns
INFO  07:45:34 Initializing system_schema.views
INFO  07:45:34 Initializing system_schema.types
INFO  07:45:34 Initializing system_schema.functions
INFO  07:45:34 Initializing system_schema.aggregates
INFO  07:45:34 Initializing system_schema.indexes
Exception (java.lang.IllegalStateException) encountered during startup: One row required, 2 found
java.lang.IllegalStateException: One row required, 2 found
    at org.apache.cassandra.cql3.UntypedResultSet$FromResultSet.one(UntypedResultSet.java:84)
    at org.apache.cassandra.schema.SchemaKeyspace.fetchTable(SchemaKeyspace.java:948)
    at org.apache.cassandra.schema.SchemaKeyspace.fetchTables(SchemaKeyspace.java:938)
    at org.apache.cassandra.schema.SchemaKeyspace.fetchKeyspace(SchemaKeyspace.java:901)
    at org.apache.cassandra.schema.SchemaKeyspace.fetchKeyspacesWithout(SchemaKeyspace.java:878)
    at org.apache.cassandra.schema.SchemaKeyspace.fetchNonSystemKeyspaces(SchemaKeyspace.java:866)
    at org.apache.cassandra.config.Schema.loadFromDisk(Schema.java:134)
    at org.apache.cassandra.config.Schema.loadFromDisk(Schema.java:124)
    at org.apache.cassandra.service.CassandraDaemon.setup(CassandraDaemon.java:229)
    at org.apache.cassandra.service.CassandraDaemon.activate(CassandraDaemon.java:551)
    at org.apache.cassandra.service.CassandraDaemon.main(CassandraDaemon.java:679)

对集群及system_schema执行修复后错误仍存在,已通过健康节点执行nodetool removenode将该节点移除,但同一DC的另一节点重启时出现相同启动失败错误。

二、CQL连接错误

无法从健康节点登录cqlsh shell,报错:

Connection error: ('Unable to connect to any servers', {'<<VM hostname>>': UnicodeDecodeError('utf8', '\x7f\x00\x00\x80C\x02', 3, 4, 'invalid start byte')})

部分节点也出现类似连接错误:

Connection error: ('Unable to connect to any servers', {'<<VM Hostname>>': ConnectionShutdown("'utf8' codec can't decode byte 0x80 in position 3: invalid start byte",)})

三、当前集群状态

仅nodetool命令可用,执行nodetool describecluster显示存在5种不同Schema版本,且有9个节点不可达:

./nodetool describecluster
Cluster Information:
    Name: Dummy cluster
    Snitch: org.apache.cassandra.locator.DynamicEndpointSnitch
    Partitioner: org.apache.cassandra.dht.Murmur3Partitioner
    Schema versions:
        1590ea6a-8c19-342a-8269-204c64a12176: [9 nodes here]
        668d9efd-13c1-3fb3-9b89-7fc07d9ddf0b: [1 node here]
        d20dc0de-dd34-3183-b459-31e3feb8f118: [3 nodes here]
        3ec9610c-d241-3215-84f2-2413b8cad7d2: [7 nodes here]
        59adb24e-f3cd-3e02-97f0-5b395827453f: [1 node here]
        UNREACHABLE: [9 nodes unreachable]

尝试在cassandra-env.sh/jvm.options中添加忽略Schema不匹配的参数,仍无法启动节点。

四、问题原因分析

  1. 系统表数据不一致:启动时的One row required, 2 found异常,说明system_schema表中存在重复元数据记录,导致Cassandra加载Schema时无法确定唯一的表结构。这种情况通常源于滚动重启过程中Schema同步异常,或之前的Schema操作(如建表/改表)未完全同步到所有节点。
  2. Schema版本碎片化:集群存在5种不同Schema版本,说明节点间Schema同步机制完全失效,各节点持有不一致的Schema信息,导致CQL连接时出现编码解码错误(不同Schema版本的元数据序列化格式不兼容)。
  3. 滚动重启中的状态破坏:第一个节点重启正常,但第二个节点出现问题,可能是第一个节点重启后打破了集群Schema同步状态,后续节点重启时无法正确拉取或加载一致的Schema。

五、恢复步骤

步骤1:确定权威Schema版本

从nodetool describecluster的Schema版本中,选择节点数量最多的版本(此处为1590ea6a-8c19-342a-8269-204c64a12176,对应9个节点)作为权威Schema版本。

步骤2:对齐健康节点到权威Schema版本

在持有权威Schema的节点上,执行以下命令强制同步Schema到其他健康节点:

nodetool refresh --full

对每个非权威版本的健康节点,依次执行:

nodetool drain
nodetool stop
# 启动节点时添加JVM参数强制拉取Schema
cassandra -Dcassandra.ignore_dc=true -Dcassandra.ignore_schema_disagreements=true

启动后等待节点加入集群,再执行nodetool describecluster确认Schema版本是否统一。

步骤3:修复无法启动的节点

对于无法启动的节点,按以下操作:

  1. 停止节点服务,备份data/system_schema目录下的所有数据。
  2. 删除data/system_schema目录下的所有文件。
  3. 修改cassandra.yaml,将auto_bootstrap设置为true(若之前为false)。
  4. 启动节点时添加强制Schema同步参数:
cassandra -Dcassandra.ignore_schema_disagreements=true -Dcassandra.join_ring=true

节点启动后会从集群拉取最新的权威Schema,自动重建system_schema表。

步骤4:处理不可达节点

对于9个不可达节点:

  1. 检查节点网络状态,确保节点间通信端口(7000、7001、9042)开放。
  2. 若节点因Schema问题无法启动,按照步骤3的方法修复。
  3. 若节点已损坏无法修复,执行nodetool removenode将其移除,之后重新加入集群。

步骤5:验证集群一致性

所有节点恢复后,执行以下命令验证:

  • nodetool describecluster:确认所有节点Schema版本一致,无不可达节点。
  • nodetool status:检查所有节点状态为UP/NORMAL。
  • 登录cqlsh,执行SELECT * FROM system_schema.keyspaces;等简单查询确认正常。

步骤6:后续预防措施

  • 滚动重启前,执行nodetool describecluster确认所有节点Schema版本完全一致。
  • 调整JVM参数时,先在单个节点测试,确认重启正常后再批量操作。
  • 定期执行nodetool repair system_schema,确保系统表数据一致性。
  • 考虑升级到更高版本的Cassandra(3.0.x版本较老,后续版本修复了大量Schema同步相关的bug)。

内容的提问来源于stack exchange,提问作者Nikhit Nair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 06:55:26