You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

给supervisord管理的进程日志加时间戳后多进程无法正常停止的问题

问题:Supervisord停止进程时主进程残留的解决办法

问题场景

为supervisord管理的alert_mgr进程日志添加时间戳,编写了bash包装脚本并配置了supervisord conf文件,时间戳功能正常,但出现异常:

  • supervisord启动后生成4个进程:1个alert_mgr主进程 + 3个bash子进程
  • 执行supervisorctl stop alert_mgr时,仅bash子进程被停止,alert_mgr主进程仍在运行

现有配置与脚本

包装脚本 wrap_timestamp.sh

#!/bin/bash

# set piped processes to be killed when script is killed
set -o pipefail

# Function to generate a timestamp with milliseconds
timestamp() {
    echo -n "$(date "+%Y-%m-%d %H:%M:%S.$(date +%N | cut -c1-3)")"
}

# Run the binary, prepend timestamps to stdout and stderr
exec "$@" 2> >(while IFS= read -r line; do echo "$(timestamp) $line" >&2; done) \
                  | while IFS= read -r line; do echo "$(timestamp) $line"; done

Supervisord配置文件

[program:alert_mgr]
command = %(ENV_bindir)s/wrap_timestamp.sh %(ENV_python3exec)s %(ENV_pythonflags)s -m cnm.services.alert_mgr
environment=MALLOC_ARENA_MAX=4
autostart = true
autorestart = true
priority = 300
oom_score_adj = -500
stopwaitsecs = 2
stdout_logfile=%(ENV_logdir)s/alert_mgr_stdout.log
stdout_logfile_maxbytes=100MB
stdout_logfile_backups=0
stderr_logfile=%(ENV_logdir)s/alert_mgr_stderr.log
stderr_logfile_maxbytes=100MB
stderr_logfile_backups=0

进程状态

ps aux | grep alert
redisla+ 57104  0.0  0.0   3848  2592 ?        S    23:39   0:00 /bin/bash /opt/redislabs/bin/wrap_timestamp.sh /opt/redislabs/bin/python3 -O -m cnm.services.alert_mgr
redisla+ 57105  0.5  1.0 240952 81304 ?        Sl   23:39   0:03 alert_mgr
redisla+ 57106  0.0  0.0   3848  1324 ?        S    23:39   0:00 /bin/bash /opt/redislabs/bin/wrap_timestamp.sh /opt/redislabs/bin/python3 -O -m cnm.services.alert_mgr
redisla+ 57107  0.0  0.0   3848  1644 ?        S    23:39   0:00 /bin/bash /opt/redislabs/bin/wrap_timestamp.sh /opt/redislabs/bin/python3 -O -m cnm.services.alert_mgr

解决方案

方法1:修改Supervisord配置(推荐)

在[program:alert_mgr]段添加以下两个参数,让supervisord停止整个进程组而非仅父进程:

stopasgroup=true
killasgroup=true
  • stopasgroup=true:向进程组发送停止信号,而非仅主进程
  • killasgroup=true:如果进程未及时响应停止信号,会向整个进程组发送kill信号

修改后的完整配置:

[program:alert_mgr]
command = %(ENV_bindir)s/wrap_timestamp.sh %(ENV_python3exec)s %(ENV_pythonflags)s -m cnm.services.alert_mgr
environment=MALLOC_ARENA_MAX=4
autostart = true
autorestart = true
priority = 300
oom_score_adj = -500
stopwaitsecs = 2
stopasgroup=true
killasgroup=true
stdout_logfile=%(ENV_logdir)s/alert_mgr_stdout.log
stdout_logfile_maxbytes=100MB
stdout_logfile_backups=0
stderr_logfile=%(ENV_logdir)s/alert_mgr_stderr.log
stderr_logfile_maxbytes=100MB
stderr_logfile_backups=0

方法2:优化包装脚本(避免子shell进程组问题)

原脚本的管道和进程替换创建了多个独立子shell,导致信号无法传递到alert_mgr主进程。可以修改脚本,让主进程成为会话组长,并确保信号能正确传递:

#!/bin/bash
set -o errexit
set -o pipefail

timestamp() {
    echo -n "$(date "+%Y-%m-%d %H:%M:%S.$(date +%N | cut -c1-3)")"
}

# 启动目标进程,将stdout和stderr合并后统一添加时间戳
"$@" 2>&1 | while IFS= read -r line; do
    # 简单区分日志输出方向,可根据实际需求调整判断逻辑
    if [[ "$line" == *"ERROR"* || "$line" == *"WARN"* ]]; then
        echo "$(timestamp) $line" >&2
    else
        echo "$(timestamp) $line"
    fi
done &

# 获取目标进程PID
TARGET_PID=$!

# 捕获停止信号,传递给目标进程
trap "kill -TERM $TARGET_PID" INT TERM

# 等待目标进程结束
wait $TARGET_PID

注:若目标进程会fork子进程,需调整PID捕获逻辑以确保覆盖所有相关进程。

原理说明

原问题核心是:bash包装脚本通过管道和进程替换创建了多个子shell,这些子shell与alert_mgr主进程分属不同进程组。默认情况下,supervisord仅向启动的bash父进程发送停止信号,alert_mgr主进程无法收到信号导致残留。通过开启stopasgroup和killasgroup,supervisord会将信号发送到整个进程组,确保所有相关进程都被停止。

内容的提问来源于stack exchange,提问作者Meni Katz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 07:32:07