Aeron Cluster交付保障疑问:节点故障下消息丢失如何处理?
Aeron Cluster节点故障时的消息交付保障问题
有评论指出:
Cluster通过仲裁协议进一步提升能力,可在节点故障时防止消息丢失。
但实际测试单节点故障场景时,仍存在消息丢失情况。测试使用Aeron代码库1.38.1版本中的io.aeron.samples.cluster.tutorial.BasicAuctionClusterClient,并对该类做了小幅修改以验证消息接收:
public void onSessionMessage( final ClientSession session, final long timestamp, final DirectBuffer buffer, final int offset, final int length, final Header header) { final long correlationId = buffer.getLong(offset + CORRELATION_ID_OFFSET); System.out.println("Received message with correlation ID " + correlationId); // 新增打印逻辑 // 其余代码保持不变 }
启动3节点集群后,1个节点被选为LEADER,随后启动BasicAuctionClusterClient向集群发送请求。当停止Leader节点后,新Leader如期选举产生,但从Leader停止到新Leader选举完成期间发送的消息从未到达集群,日志中的correlation ID存在明显间隙:
New role is LEADER Received message with correlation ID -8046281870845246166 attemptBid(this=Auction{bestPrice=144, currentWinningCustomerId=1}, price=152,customerId=1) Received message with correlation ID -8046281870845246165 attemptBid(this=Auction{bestPrice=152, currentWinningCustomerId=1}, price=158,customerId=1) Consensus Module io.aeron.cluster.client.ClusterEvent: WARN - leader heartbeat timeout Received message with correlation ID -8046281870845246154 attemptBid(this=Auction{bestPrice=158, currentWinningCustomerId=1}, price=167,customerId=1)
若要实现交付(处理)保障,开发者需采取什么措施?是否需要自定义ACK系统、重试机制及集群节点侧的重复请求处理逻辑?
内容的提问来源于stack exchange,提问作者Frank
相关产品推荐
相关产品推荐

