ActiveMQ Artemis客户端故障转移未优先连接对应备份代理的问题求助
ActiveMQ Artemis客户端故障转移未优先连接对应备份代理的问题求助
我正在测试ActiveMQ Artemis集群的客户端故障转移功能,但实际行为和预期不符,想请大家帮忙排查下问题。
问题现象
我原本预期:当某个主节点故障时,连接到该节点的客户端应该自动切换到对应的备份节点。但实际观察到的情况是,客户端会直接跳转到集群里的其他主节点,只有偶尔会按预期连接到备份节点,而且看起来和消费者初始创建的位置有关。
举个具体例子:
- 消费者A初始连接到代理A,消费者B初始连接到代理B
- 关闭代理A后,消费者A并没有切换到代理A的备份节点,反而直接连到了代理B
- 重启代理A后再关闭代理B,之前跳转到代理B的消费者会切回代理A,而另一个消费者则正确连接到代理B的备份节点
环境说明
我用的是复制式HA架构,目前在本地部署了6个节点(3主3备,通过端口区分),后续计划部署到3台物理服务器,现在先做本地测试验证。
客户端配置
连接工厂代码
我参考了官方的客户端故障转移相关文档,在连接URL里添加了ha=true和reconnectAttempts=10参数,代码如下:
public ActiveMQConnectionFactory activeMQConnectionFactory() throws NamingException{ ArtemisTargetConfig config = (ArtemisTargetConfig) new InitialContext().lookup("java:comp/env/bean/ArtemisConfig"); ActiveMQConnectionFactory cf = new ActiveMQConnectionFactory("udp://" + config.getGroupAddress() + ":" + config.getGroupPort() + "?ha=true&reconnectAttempts=10"); cf.setUser(config.getUserName()); cf.setPassword(config.getPassword()); return cf; }
最终生成的连接URL是:udp://231.7.7.7:9876?ha=true&reconnectAttempts=10
JmsConfig配置
@EnableJms @Configuration public class JmsConfig { private static final Logger LOG = LoggerFactory.getLogger(JmsConfig.class); private final ActiveMQConnectionFactory activeMQConnectionFactory; @Autowired public JmsConfig( ActiveMQConnectionFactory activeMQConnectionFactory){ this.activeMQConnectionFactory = activeMQConnectionFactory; } @Bean public CachingConnectionFactory cachingConnectionFactory() { return new CachingConnectionFactory(activeMQConnectionFactory); } @Bean public JmsTemplate jmsTemplate(CachingConnectionFactory cachingConnectionFactory) { JmsTemplate template = new JmsTemplate(cachingConnectionFactory); template.setPubSubDomain(false); return template; } @Bean public DefaultJmsListenerContainerFactory jmsListenerContainerFactory () { DefaultJmsListenerContainerFactory jmsListenerContainerFactory = new DefaultJmsListenerContainerFactory(); jmsListenerContainerFactory.setConnectionFactory(activeMQConnectionFactory); jmsListenerContainerFactory.setSessionAcknowledgeMode(ActiveMQJMSConstants.INDIVIDUAL_ACKNOWLEDGE); jmsListenerContainerFactory.setPubSubDomain(false); MappingJackson2MessageConverter converter = new MappingJackson2MessageConverter(); ObjectMapper objectMapper = new ObjectMapper(); objectMapper.registerModule(new JavaTimeModule()); converter.setObjectMapper(objectMapper); converter.setTargetType(MessageType.TEXT); converter.setTypeIdPropertyName("_type"); jmsListenerContainerFactory.setCacheLevel(null); jmsListenerContainerFactory.setMessageConverter(converter); return jmsListenerContainerFactory; } }
代理配置(关键部分)
主节点1的broker.xml
<name>localhost</name> <connectors> <connector name="artemis">tcp://localhost:61616</connector> </connectors> <acceptors> <acceptor name="artemis">tcp://localhost:61616?tcpSendBufferSize=1048576;tcpReceiveBufferSize=1048576;amqpMinLargeMessageSize=102400;protocols=CORE,AMQP,STOMP,HORNETQ,MQTT,OPENWIRE;useEpoll=true;amqpCredits=1000;amqpLowCredits=300;amqpDuplicateDetection=true;supportAdvisory=false;suppressInternalManagementObjects=false </acceptor> </acceptors> <cluster-user>cluster-admin</cluster-user> <cluster-password>apassword</cluster-password> <broadcast-groups> <broadcast-group name="bg-group1"> <group-address>231.7.7.7</group-address> <group-port>9876</group-port> <broadcast-period>5000</broadcast-period> <connector-ref>artemis</connector-ref> </broadcast-group> </broadcast-groups> <discovery-groups> <discovery-group name="dg-group1"> <group-address>231.7.7.7</group-address> <group-port>9876</group-port> <refresh-timeout>10000</refresh-timeout> </discovery-group> </discovery-groups> <cluster-connections> <cluster-connection name="my-cluster"> <connector-ref>artemis</connector-ref> <message-load-balancing>ON_DEMAND</message-load-balancing> <max-hops>1</max-hops> <discovery-group-ref discovery-group-name="dg-group1"/> </cluster-connection> </cluster-connections> <ha-policy> <replication> <primary> <vote-on-replication-failure>true</vote-on-replication-failure> <check-for-active-server>true</check-for-active-server> </primary> </replication> </ha-policy>
备份节点1的broker.xml
<name>localhost</name> <connectors> <connector name="artemis">tcp://localhost:61617</connector> </connectors> <acceptors> <!-- Acceptor for every supported protocol --> <acceptor name="artemis">tcp://localhost:61617?tcpSendBufferSize=1048576;tcpReceiveBufferSize=1048576;amqpMinLargeMessageSize=102400;protocols=CORE,AMQP,STOMP,HORNETQ,MQTT,OPENWIRE;useEpoll=true;amqpCredits=1000;amqpLowCredits=300;amqpDuplicateDetection=true;supportAdvisory=false;suppressInternalManagementObjects=false </acceptor> </acceptors> <cluster-user>cluster-admin</cluster-user> <cluster-password>aPassword</cluster-password> <broadcast-groups> <broadcast-group name="bg-group1"> <group-address>231.7.7.7</group-address> <group-port>9876</group-port> <broadcast-period>5000</broadcast-period> <connector-ref>artemis</connector-ref> </broadcast-group> </broadcast-groups> <discovery-groups> <discovery-group name="dg-group1"> <group-address>231.7.7.7</group-address> <group-port>9876</group-port> <refresh-timeout>10000</refresh-timeout> </discovery-group> </discovery-groups> <cluster-connections> <cluster-connection name="my-cluster"> <connector-ref>artemis</connector-ref> <message-load-balancing>ON_DEMAND</message-load-balancing> <max-hops>1</max-hops> <discovery-group-ref discovery-group-name="dg-group1"/> </cluster-connection> </cluster-connections> <ha-policy> <replication> <backup> <allow-failback>true</allow-failback> </backup> </replication> </ha-policy> </core>
集群拓扑信息
通过代理的listNetworkTopology操作得到的结果:
[{"nodeID":"19f26f12-ee1f-11ef-b81c-00155df5022d","live":"localhost:61616","primary":"localhost:61616","backup":"localhost:61617"}, {"nodeID":"b071e7f0-ee3f-11ef-b10d-00155df5022d","live":"localhost:61620","primary":"localhost:61620","backup":"localhost:61621"}, {"nodeID":"d46e2693-ee1f-11ef-9cd3-00155df5022d","live":"localhost:61618","primary":"localhost:61618","backup":"localhost:61619"}]
已尝试的排查操作
- 确认复制功能正常:主节点日志能看到
AMQ221025: Replication: sending NIOSequentialFile,备份节点日志有AMQ221031: backup announced的输出 - 完全移除了
CachedConnectionFactory,问题依旧 - 删除并重建过所有代理,问题短暂解决但之后又重现
有没有朋友遇到过类似的情况?或者能帮我指出配置里可能存在的问题?
备注:内容来源于stack exchange,提问作者Sergio Scaramuzzi
相关产品推荐
相关产品推荐

