ActiveMQ主备HA选主Python脚本异常:服务化或Cron方案咨询
Python脚本功能说明
为Artemis ActiveMQ主备高可用(HA)实现了leader election(选主机制)。该脚本会检测ActiveMQ Pod的可用性,并更新Pod的标签(Service的标签固定为primary),从而帮助Service将流量路由到当前的主Pod。
伪代码:
while true: if currentPod's TCP socket 61616 is running: change the label of the pod to 'primary' if not already 'primary' else: change the label of the pod to 'backup' if not already 'backup' wait(3 sec)
Pod基础镜像:eclipse-temurin:17
当前该脚本运行正常,当主Pod崩溃或重启时,流量会顺利切换到备用Pod。
当前运行方式
在entrypoint.sh中通过以下命令启动:
python3 $PATH/leader-election.py &
Dockerfile与entrypoint.sh相关代码
以下是与本问题相关的核心代码片段:
Dockerfile:
# Create broker instance FROM eclipse-temurin:17 as builder ARG ACTIVEMQ_ARTEMIS_VERSION ENV ACTIVEMQ_ARTEMIS_VERSION=$ACTIVEMQ_ARTEMIS_VERSION RUN apt-get -qq -o=Dpkg::Use-Pty=0 update && apt-get -y install && apt-get -y install wget RUN apt-get install -y python3 RUN apt-get install -y python3-pip RUN wget https://archive.apache.org/dist/activemq/activemq-artemis/$ACTIVEMQ_ARTEMIS_VERSION/apache-artemis-$ACTIVEMQ_ARTEMIS_VERSION-bin.tar.gz RUN tar xvfz apache-artemis-$ACTIVEMQ_ARTEMIS_VERSION-bin.tar.gz RUN ln -s /opt/apache-artemis-${ACTIVEMQ_ARTEMIS_VERSION} /opt/apache-artemis RUN rm -f apache-artemis-${ACTIVEMQ_ARTEMIS_VERSION}-bin.tar.gz KEYS apache-artemis-${ACTIVEMQ_ARTEMIS_VERSION}-bin.tar.gz.asc WORKDIR /var/lib RUN "/opt/apache-artemis-${ACTIVEMQ_ARTEMIS_VERSION}/bin/artemis" create artemis \ --home /opt/apache-artemis \ --http-host 0.0.0.0 \ --user artemis \ --password artemis \ --no-hornetq-acceptor --no-mqtt-acceptor --no-stomp-acceptor --no-amqp-acceptor \ --role amq \ --require-login ; # some more configs COPY assets/docker-entrypoint.sh / ENTRYPOINT ["/docker-entrypoint.sh"] CMD ["artemis-server"]
entrypoint.sh:
if [ "$LEADER_ELECTION" = "true" ]; then python3 $OVERRIDE_PATH/leader-election.py & fi # some more things if [ "$1" = 'artemis-server' ]; then exec dumb-init -- sh ./artemis run fi
StatefulSet与Service配置
Service配置:
apiVersion: v1 kind: Service metadata: name: activemq-svc namespace: activemq spec: type: ClusterIP ports: - port: 61616 name: netty-connector protocol: TCP targetPort: 61616 selector: app: activemq node: primary
StatefulSet配置:
apiVersion: apps/v1 kind: StatefulSet metadata: name: activemq-statefulset namespace: activemq spec: replicas: 2 serviceName: activemq-svc selector: matchLabels: app: activemq template: metadata: labels: app: activemq node: not-ready
Python脚本会根据Pod可用性将node: not-ready标签更新为node: primary或node: backup,确保同一时刻仅有一个Pod为primary,使Service仅将流量路由至该Pod。
问题解答
1. 如何将Python脚本以服务方式运行?
当前脚本后台异步启动(&)无进程守护,易意外退出,可通过以下方式做成可靠服务:
- 用supervisord托管:适合容器环境,无需systemd。先安装supervisord:
apt-get install -y supervisor,创建配置文件/etc/supervisor/conf.d/leader-election.conf:
然后在entrypoint.sh中启动supervisord,同时管理ActiveMQ进程。[program:leader-election] command=/usr/bin/python3 $OVERRIDE_PATH/leader-election.py autostart=true autorestart=true stderr_logfile=/var/log/leader-election.err.log stdout_logfile=/var/log/leader-election.out.log - 用systemd管理:容器内创建
/etc/systemd/system/leader-election.service:
调整容器启动命令为systemd,在entrypoint.sh中执行[Unit] Description=ActiveMQ Leader Election Script After=network.target [Service] ExecStart=/usr/bin/python3 $OVERRIDE_PATH/leader-election.py Restart=always RestartSec=3 User=root [Install] WantedBy=multi-user.targetsystemctl enable --now leader-election.service。 - 添加简易重启逻辑:修改entrypoint.sh,用bash循环保证脚本退出后自动重启:
if [ "$LEADER_ELECTION" = "true" ]; then while true; do python3 $OVERRIDE_PATH/leader-election.py echo "Leader election script exited, restarting in 3s..." sleep 3 done & fi
2. 是否可通过Cron每3秒执行脚本?
Cron最小执行间隔为1分钟,无法直接实现每3秒执行。若用定时方式,只能用bash循环+sleep模拟,但每次执行都要重新初始化(如连接K8s API),效率低于长进程,且无法及时检测间隔内的Pod状态变化,不推荐。
3. 能否不使用无限循环实现选主逻辑?
可以,推荐以下替代方案:
- 用Kubernetes原生选主机制:借助K8s官方的leader election客户端库(支持Python),基于ConfigMap或Lease资源实现分布式锁,由集群协调选主,可靠性更高,主节点故障时备用节点自动抢占锁并更新标签。
- 利用Artemis原生HA:配置Artemis主备集群(共享存储或复制模式),让Broker自身管理主备状态,结合K8s就绪探针检测Broker主状态,动态更新Pod标签或调整Service路由。
- 外部控制器处理:用K8s Operator或自定义Controller监听Pod的就绪/存活状态,自动更新标签,将检测与标签更新解耦,避免在Pod内运行脚本。
内容的提问来源于stack exchange,提问作者Subhidh Agarwal
相关产品推荐
相关产品推荐

