You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

异步监控并中断Shell脚本的问题咨询及优化建议

问题分析与解决方案

1. 避免cleanup函数被调用两次的问题

你的cleanup被触发两次的核心原因是trap cleanup SIGTERM EXIT的设置:当monitor_status发送kill -15 $$时,主进程先收到SIGTERM信号触发一次cleanup;随后主进程因终止信号退出,又会触发EXIT信号的trap,导致cleanup重复执行。

推荐两种解决方式:

方式一:用标记变量防止重复执行

在cleanup中加入全局标记,确保仅执行一次:

#!/bin/bash
CLEANUP_DONE=0

function cleanup() {
    if [[ $CLEANUP_DONE -eq 1 ]]; then
        return
    fi
    CLEANUP_DONE=1
    echo "cleanup invoked"
    if [[ -n "${child}" ]]; then
        echo "Stopping the child"
        kill "${child}" >/dev/null 2>&1 || true
    fi
    if [[ -n "${monitor_pid}" ]]; then
        echo "Stopping monitor"
        kill "${monitor_pid}" >/dev/null 2>&1 || true
    fi
    echo "cleanup done"
}

方式二:分离SIGTERM与EXIT的处理逻辑

如果两种场景的清理需求有差异,可以拆分trap逻辑:

# 处理强制终止信号
trap 'echo "Received SIGTERM, initiating cleanup..."; cleanup_sigterm; exit 1' SIGTERM
# 处理正常退出
trap cleanup_exit EXIT

function cleanup_sigterm() {
    # 终止子进程与监控进程
    [[ -n "${child}" ]] && kill "${child}" >/dev/null 2>&1 || true
    [[ -n "${monitor_pid}" ]] && kill "${monitor_pid}" >/dev/null 2>&1 || true
}

function cleanup_exit() {
    echo "Performing normal exit cleanup"
    # 正常退出的专属清理逻辑
}

2. 主进程正常完成时终止monitor_status

你需要先保存monitor_status的PID,在主进程结束后主动终止它:

# 启动监控进程并保存PID
monitor_status &
monitor_pid=$!

# 启动主进程并保存PID
main &
child=$!

# 等待主进程完成
wait "${child}"

# 主进程正常结束后,终止监控进程
kill "${monitor_pid}" >/dev/null 2>&1 || true

另一种更健壮的方式是让monitor_status主动检查主进程是否存活:

function monitor_status() {
    while true; do
        echo "checking status"
        # 先检查主进程是否还在运行
        if ! kill -0 "${child}" >/dev/null 2>&1; then
            echo "Main process exited, stopping monitor"
            exit 0
        fi
        # 修正语法错误:原代码的赋值方式不正确
        Status=$(get_status)
        if [[ "${Status}" == "Terminate" ]]; then
            echo "Alert Alert!!!"
            kill -15 $$
        fi
        sleep 5
    done
}

注:原脚本的monitor_status存在两处语法错误:缺少while true; do循环(导致监控只执行一次就退出),赋值语句$Status = get_status应改为Status=$(get_status)。

3. 更佳方案与脚本改进建议

更佳方案:使用进程组统一管理

将主进程与监控进程归入同一个进程组,这样发送信号时可以一次性终止所有相关进程,避免遗漏:

#!/bin/bash
set -m # 启用作业控制

CLEANUP_DONE=0

function cleanup() {
    if [[ $CLEANUP_DONE -eq 1 ]]; then
        return
    fi
    CLEANUP_DONE=1
    echo "cleanup invoked"
    # 终止整个进程组(包含主进程、监控进程)
    kill -- -$$ >/dev/null 2>&1 || true
    echo "cleanup done"
}

function main() {
    echo "Starting main process"
    # 模拟长时间数据处理
    sleep 30
    echo "Exiting Main"
}

function monitor_status() {
    while true; do
        echo "checking status"
        # 替换为你的实际状态检测逻辑
        Status=$(get_status)
        if [[ "${Status}" == "Terminate" ]]; then
            echo "Alert Alert!!!"
            kill -15 -- -$$ # 终止整个进程组
            exit 0
        fi
        sleep 5
    done
}

trap cleanup SIGTERM EXIT

# 启动两个进程,自动加入同一进程组
monitor_status &
main &

wait

其他改进点

  • 开启严格模式:在脚本开头加入set -euo pipefail,捕获更多潜在错误,提升脚本健壮性。
  • 日志持久化:将echo输出替换为写入日志文件,例如echo "cleanup invoked" >> /var/log/data_processing.log,方便后续排查问题。
  • 资源预检查:在执行get_status前,先验证其依赖的外部资源(如API、文件)是否可用,避免无意义的循环等待。
  • 关键步骤信号屏蔽:在执行不可中断的数据处理步骤时,临时屏蔽SIGTERM信号,防止数据损坏:
    function main() {
        echo "Starting main process"
        # 屏蔽SIGTERM
        trap '' SIGTERM
        # 执行关键数据处理
        process_large_dataset
        # 恢复SIGTERM处理
        trap cleanup SIGTERM
        echo "Exiting Main"
    }
    
  • PID文件管理:如果脚本需要后台运行,创建PID文件记录主进程PID,防止重复执行。

内容的提问来源于stack exchange,提问作者eeedev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 18:57:40