You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Aeron Cluster交付保障疑问:节点故障下消息丢失如何处理?

Aeron Cluster节点故障时的消息交付保障问题

有评论指出:

Cluster通过仲裁协议进一步提升能力,可在节点故障时防止消息丢失。

但实际测试单节点故障场景时,仍存在消息丢失情况。测试使用Aeron代码库1.38.1版本中的io.aeron.samples.cluster.tutorial.BasicAuctionClusterClient,并对该类做了小幅修改以验证消息接收:

public void onSessionMessage(
    final ClientSession session,
    final long timestamp,
    final DirectBuffer buffer,
    final int offset,
    final int length,
    final Header header)
{
    final long correlationId = buffer.getLong(offset + CORRELATION_ID_OFFSET);                   
    System.out.println("Received message with correlation ID " + correlationId); // 新增打印逻辑
    // 其余代码保持不变
}

启动3节点集群后,1个节点被选为LEADER,随后启动BasicAuctionClusterClient向集群发送请求。当停止Leader节点后,新Leader如期选举产生,但从Leader停止到新Leader选举完成期间发送的消息从未到达集群,日志中的correlation ID存在明显间隙:

New role is LEADER
Received message with correlation ID -8046281870845246166
attemptBid(this=Auction{bestPrice=144, currentWinningCustomerId=1}, price=152,customerId=1)
Received message with correlation ID -8046281870845246165
attemptBid(this=Auction{bestPrice=152, currentWinningCustomerId=1}, price=158,customerId=1)
Consensus Module
io.aeron.cluster.client.ClusterEvent: WARN - leader heartbeat timeout
Received message with correlation ID -8046281870845246154
attemptBid(this=Auction{bestPrice=158, currentWinningCustomerId=1}, price=167,customerId=1)

若要实现交付(处理)保障,开发者需采取什么措施?是否需要自定义ACK系统、重试机制及集群节点侧的重复请求处理逻辑?

内容的提问来源于stack exchange,提问作者Frank

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 16:24:31