Linux服务器中Python监控脚本如何在新Shell重启崩溃的Python脚本
问题描述
我的Linux服务器开机时会运行4个Python脚本,但有时部分脚本会崩溃,需要手动重启。因此我编写了一个监控脚本,通过ps -fA命令检测这些脚本是否停止运行,若停止则重启,代码如下:
# Search output generaldata = "site_general_data" amsniffer = "site_AM_sniffer" b3sniffer = "site_B3_sniffer" cryptosniffer = "site_CRYPTO_snifferB" # Output active outputa = subprocess.check_output('ps -fA | grep python',shell=True) # Verify Scripts if generaldata not in str(outputa): subprocess.run('sudo python3.6 /bin/site_general_data.py', shell=True) ativogen = "General Data inactive" print("General Data inactive") else: ativogen = "General Data active" if amsniffer not in str(outputa): subprocess.run('sudo python3.6 /bin/site_AM_sniffer.py', shell=True) ativoam = "AM Sniffer inactive" print("AM Sniffer inactive") else: ativoam = "AM Sniffer active" if b3sniffer not in str(outputa): subprocess.run('sudo python3.6 /bin/site_B3_sniffer.py', shell=True) ativob3 = "B3 Sniffer inactive" print("B3 Sniffer inactive") else: ativob3 = "B3 Sniffer active" if cryptosniffer not in str(outputa): subprocess.run('sudo python3.6 /bin/site_CRYPTO_snifferB.py', shell=True) ativocry = "Crypto Sniffer inactive" print("Crypto Sniffer inactive") else: ativocry = "Crypto Sniffer active"
涉及的脚本为site_general_data、site_AM_sniffer、site_B3_sniffer和site_CRYPTO_snifferB。当前问题是:我希望监控脚本持续运行不中断,重启崩溃的脚本时要在新Shell中启动,但现在重启的脚本会在同一个Shell中运行。已查阅subprocess模块的Popen文档,但仍未解决问题,请问有实现方法吗?
解决方案
方法一:用nohup结合后台运行
在启动脚本的命令里加入nohup和后台符号&,让脚本脱离当前Shell独立运行,同时重定向输出避免占用终端:
修改subprocess.run的命令为:
subprocess.run('sudo nohup python3.6 /bin/site_general_data.py > /dev/null 2>&1 &', shell=True)
nohup:保证进程在当前Shell退出后仍能继续运行> /dev/null 2>&1:把标准输出和错误输出重定向到空设备,避免生成冗余的nohup.out文件&:让进程在后台运行,不阻塞监控脚本
方法二:用subprocess.Popen开启新会话
直接替换subprocess.run为Popen,并设置start_new_session=True,让子进程脱离当前Shell的控制:
import subprocess # 以site_general_data为例 if generaldata not in str(outputa): subprocess.Popen(['sudo', 'python3.6', '/bin/site_general_data.py'], start_new_session=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) ativogen = "General Data inactive" print("General Data inactive")
start_new_session=True:创建新的会话组,子进程不会随监控脚本的Shell退出而终止stdout和stderr重定向到DEVNULL,避免脚本输出干扰监控逻辑
方法三:配置成systemd系统服务(推荐长期使用)
如果需要稳定的长期监控,更推荐把这4个脚本做成systemd服务,利用系统自带的重启机制,无需自己维护监控脚本:
- 为每个脚本创建服务文件,比如
/etc/systemd/system/site-general-data.service:
[Unit] Description=Site General Data Script After=network.target [Service] Type=simple User=root ExecStart=/usr/bin/python3.6 /bin/site_general_data.py Restart=always RestartSec=5 [Install] WantedBy=multi-user.target
- 启用并启动服务:
sudo systemctl daemon-reload sudo systemctl enable site-general-data.service sudo systemctl start site-general-data.service
Restart=always:进程无论因何种原因退出,都会自动重启RestartSec=5:重启前等待5秒,避免进程崩溃后频繁重启
优化现有监控的检测逻辑
当前用ps -fA | grep python的方式容易误判(比如grep进程本身会被匹配),可以改用更精准的检测方式:
def is_script_running(script_name): # 用pgrep精确匹配进程名 result = subprocess.run(['pgrep', '-f', script_name], capture_output=True, text=True) return len(result.stdout.strip()) > 0
之后调用is_script_running(generaldata)就能准确判断脚本是否在运行,避免误触发重启。
内容的提问来源于stack exchange,提问作者Capuccino
相关产品推荐
相关产品推荐

