You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes中注入Istio的ActiveMQ Artemis集群连接异常求助

根因分析
  1. Istio代理TCP策略冲突:尽管已设置idleTimeout为无限,但Envoy默认的max_connection_duration、connect_timeout等TCP参数仍可能主动断开长连接;同时Istio对集群内部TCP流量的拦截会干扰Artemis主备间的心跳机制,导致主备连接被标记为失效并发送RST包。
  2. Artemis与EAP的连接存活配置不匹配:仅配置EAP的重连参数未同步调整Artemis侧的心跳与连接超时设置,导致连接在Istio代理层面被断开后,两端未及时触发重连或心跳维持。
  3. 存活探针配置不当:若存活探针直接使用Artemis的集群通信/JMS端口,Istio代理会将探测流量视为普通连接,频繁的探测可能干扰长连接的稳定性,甚至触发连接重置。
补充配置方案

Apache ActiveMQ Artemis(broker.xml)

  • 调整集群连接器的连接与心跳参数:
    <cluster-connections>
      <cluster-connection name="my-cluster">
        <connector-ref>artemis-connector</connector-ref>
        <retry-interval>1000</retry-interval>
        <retry-interval-multiplier>2</retry-interval-multiplier>
        <max-retry-interval>60000</max-retry-interval>
        <reconnect-attempts>-1</reconnect-attempts>
        <connection-ttl>300000</connection-ttl> <!-- 5分钟,高于Istio可能的超时阈值 -->
        <call-timeout>30000</call-timeout>
        <tcp-nodelay>true</tcp-nodelay> <!-- 禁用Nagle算法,确保心跳及时发送 -->
      </cluster-connection>
    </cluster-connections>
    
  • 配置接受器的TCP优化:
    <acceptors>
      <acceptor name="core">tcp://0.0.0.0:61616?tcp-nodelay=true&amp;keepAlive=true</acceptor>
      <acceptor name="amqp">tcp://0.0.0.0:5672?tcp-nodelay=true&amp;keepAlive=true</acceptor>
    </acceptors>
    

JBoss EAP(standalone-full.xml)

  • 完善连接池的重连与验证配置:
    <pooled-connection-factory name="activemq-ra" entries="java:/JmsXA java:jboss/DefaultJMSConnectionFactory" connectors="activemq-connector" transaction="xa">
      <reconnect-attempts>-1</reconnect-attempts>
      <connection-ttl>86400000</connection-ttl>
      <retry-interval>1000</retry-interval>
      <retry-interval-multiplier>2</retry-interval-multiplier>
      <max-retry-interval>60000</max-retry-interval>
      <validate-on-match>true</validate-on-match> <!-- 获取连接时验证有效性 -->
      <ha>true</ha> <!-- 启用HA模式,感知主备切换 -->
    </pooled-connection-factory>
    

Istio 配置

  • 更新EnvoyFilter,覆盖所有TCP连接限制:
    apiVersion: networking.istio.io/v1alpha3
    kind: EnvoyFilter
    metadata:
      name: artemis-tcp-timeouts
      namespace: istio-system
    spec:
      workloadSelector:
        labels:
          app: artemis-cluster
      configPatches:
      - applyTo: NETWORK_FILTER
        match:
          context: SIDECAR_INBOUND
          listener:
            portNumber: 61616
            filterChain:
              filter:
                name: "envoy.filters.network.tcp_proxy"
        patch:
          operation: MERGE
          value:
            name: "envoy.filters.network.tcp_proxy"
            typed_config:
              "@type": "type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy"
              idle_timeout: 0s # 无限超时
              max_connection_duration: 0s # 禁用最大连接时长限制
      - applyTo: NETWORK_FILTER
        match:
          context: SIDECAR_OUTBOUND
          listener:
            filterChain:
              filter:
                name: "envoy.filters.network.tcp_proxy"
        patch:
          operation: MERGE
          value:
            name: "envoy.filters.network.tcp_proxy"
            typed_config:
              "@type": "type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy"
              idle_timeout: 0s
              max_connection_duration: 0s
    
  • 配置DestinationRule优化连接池:
    apiVersion: networking.istio.io/v1alpha3
    kind: DestinationRule
    metadata:
      name: artemis-service
      namespace: default
    spec:
      host: artemis-cluster.default.svc.cluster.local
      trafficPolicy:
        connectionPool:
          tcp:
            connectTimeout: 30s
            tcpKeepalive:
              time: 7200s # 2小时发送一次TCP keepalive
              interval: 75s
    
  • 调整Sidecar的流量策略:
    apiVersion: networking.istio.io/v1alpha3
    kind: Sidecar
    metadata:
      name: artemis-sidecar
      namespace: default
    spec:
      workloadSelector:
        labels:
          app: artemis-cluster
      outboundTrafficPolicy:
        mode: ALLOW_ANY # 允许所有出站流量,避免Istio拦截集群内部通信
    
存活检查确认建议
  1. 替换TCP探针为应用级健康检查:避免直接使用Artemis的61616/5672端口作为探针端口,改用Artemis的CLI命令:
    livenessProbe:
      exec:
        command: ["/opt/artemis/bin/artemis", "check", "health"]
      initialDelaySeconds: 30
      periodSeconds: 30
      timeoutSeconds: 5
      failureThreshold: 3
    readinessProbe:
      exec:
        command: ["/opt/artemis/bin/artemis", "check", "health"]
      initialDelaySeconds: 10
      periodSeconds: 15
      timeoutSeconds: 5
      failureThreshold: 3
    
  2. 避免探针干扰长连接:确保存活探针的periodSeconds不小于30秒,减少探测频率对长连接的影响。
  3. 验证Istio探针流量处理:确认Istio代理不会将探针流量计入普通连接统计,避免因探测导致连接池耗尽或重置。

内容的提问来源于stack exchange,提问作者MYC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 08:48:11