ClickHouse集群Keeper同步DDL但数据不复制,报Session expired错误
ClickHouse集群数据同步异常(Session expired)排查解决
问题现象
- Keeper进程可正常同步DDL语句,ReplicatedMergeTree类型表能在所有节点成功创建
- 执行数据插入后,仅执行插入操作的节点存在数据,其余节点无同步数据
- 远程服务器日志报错:
Session expired(KEEPER_EXCEPTION)
相关配置与操作语句
1. 集群远程配置(remote_servers)
<remote_servers replace="true"> <clustertest> <secret>key</secret> <shard> <internal_replication>true</internal_replication> <replica> <host>nod1</host> <port>3307</port> </replica> </shard> <shard> <internal_replication>true</internal_replication> <replica> <host>nod2</host> <port>3307</port> </replica> </shard> <shard> <internal_replication>true</internal_replication> <replica> <host>nod3</host> <port>3307</port> </replica> </shard> </cluster_gbit> <!-- 注意:此处标签不匹配,应为</clustertest> --> </remote_servers>
2. 宏配置(macros)
<macros> <shard>1</shard> <replica>01</replica> <cluster>clustertest</cluster> </macros>
3. Keeper配置(keeper_server)
<keeper_server> <tcp_port>19181</tcp_port> <server_id>1</server_id> <log_storage_path>/var/lib/clickhouse/coordination/logs</log_storage_path> <snapshot_storage_path>/var/lib/clickhouse/coordination/snapshots</snapshot_storage_path> <coordination_settings> <operation_timeout_ms>10000</operation_timeout_ms> <min_session_timeout_ms>10000</min_session_timeout_ms> <session_timeout_ms>100000</session_timeout_ms> <raft_logs_level>information</raft_log_level> <compress_logs>false</compress_logs> </coordination_settings> <hostname_checks_enabled>true</hostname_checks_enabled> <raft_configuration> <server> <id>1</id> <hostname>nod1</hostname> <port>19234</port> </server> <server> <id>2</id> <hostname>nod2</hostname> <port>19234</port> </server> <server> <id>3</id> <hostname>nod3</hostname> <port>19234</port> </server> <server> <id>5</id> <hostname>keeper</hostname> <port>19234</port> </server> </raft_configuration> </keeper_server>
4. 建表与插入语句
CREATE TABLE dwhcluster.table1 ON CLUSTER clustertest ( `id` UInt64, `column1` String ) ENGINE = ReplicatedMergeTree ORDER BY id INSERT INTO dwhcluster.table1 (id, column1) VALUES (3, 'abc'), (4, 'def')
5. 报错日志
ERROR : dwhcluster.table1 (ae4d40e3-b7e5-4a70-9268-20dcdcbab970): void DB::StorageReplicatedMergeTree::mutationsUpdatingTask(): Code: 999. Coordination::Exception: Session expired. (KEEPER_EXCEPTION), Stack trace (when copying this message, always include the lines below):
排查与解决思路
1. 修复配置语法错误
优先修正remote_servers中的标签不匹配问题:将</cluster_gbit>改为</clustertest>,确保集群拓扑配置能被正确解析,之后重启所有ClickHouse节点。
2. 统一Keeper会话超时参数
- 确保所有节点的Keeper配置中,
session_timeout_ms和min_session_timeout_ms参数值完全一致,避免节点间会话超时规则不匹配导致会话提前失效 - 若节点间网络延迟较高,可适当调大
session_timeout_ms(例如改为300000,即5分钟),同时保证operation_timeout_ms小于session_timeout_ms
3. 验证Keeper集群状态
执行命令检查Keeper集群节点是否正常连通:
clickhouse-keeper-client --host nod1 --port 19181
在客户端中输入stat查看集群状态,确认所有配置的Keeper节点(nod1、nod2、nod3、keeper)均处于正常在线状态。
4. 检查ZooKeeper路径权限
通过ZooKeeper客户端确认ClickHouse进程对表的ZooKeeper存储路径有读写权限:
zkCli.sh -server nod1:19181 ls /clickhouse/tables/dwhcluster/table1
若权限不足,需调整ZooKeeper的ACL配置,或确保ClickHouse进程运行用户拥有对应权限。
5. 重启异常节点
若上述配置修复后数据仍未同步,可重启未同步数据的ClickHouse节点,让节点重新与Keeper建立会话并拉取同步任务。
内容的提问来源于stack exchange,提问作者Jefferson cardona chacua
相关产品推荐
相关产品推荐

