如何用Python重置Systemd Watchdog?多线程图片检测服务重启方案
Systemd Watchdog 配置问题排查与解决方案
问题描述
我正在为多线程图片检测软件实现Systemd Watchdog,该软件依赖较多。之前通过shell脚本启动服务,现在改为直接启动Python文件,但Watchdog功能异常,服务每30秒自动重启。尝试过两种喂狗方式:写入/dev/watchdog和调用systemd.daemon.notify("WATCHDOG=1"),均未解决。
曾尝试自定义"FileWatchdog",通过shell脚本发送错误信息并重启,但该方案需要root权限,且硬编码密码不利于分发,维护成本高。
需求:当主程序陷入循环卡顿30秒及以上时,自动重启"Picture Detection Main Application"服务。同时有疑问:这是多线程程序,主进程卡住时,服务能否正常重启?
当前Systemd服务配置
[Unit] Description=Picturedetection Main application Wants=network-online.target After=network-online.target [Service] Type=simple User=user WorkingDirectory=/home/user/detection/ ExecStart=/usr/bin/python3 /home/user/detection/picturedetection.py Environment=TF_CUDNN_USE_AUTOTUNE=0 WatchdogSec=30 Restart=always WatchdogTimestamp=30 [Install] WantedBy=multi-user.target
当前Python主程序代码
import sys import syslog from multiprocessing import Queue from DetectionDefines import Detection_Version as OV import time print("OPTICONTROL START") syslog.syslog(syslog.LOG_NOTICE, "PICTUREDETECTION START --- Version " + OV.major + "." + OV.minor) from config.Config import Config as conf from prediction.ImageFeed import ImageFeed from prediction.ResultHandler import ResultHandler from dataflow.CommServer import CommServer from dataflow.FTLComm import FTLComm from dataflow.MiniHTTPServer import MiniHTTPServer from dataflow.GraphDownloader import GraphDownloader from tools.Logger import Logger from dataflow.FTPHandler import FTPHandler from tools.FileJanitor import FileJanitor from prediction.PredictionPipeline import PredictionPipeline #Watchdog test import os import time import systemd # Communication CommServer().start() FTLComm() #Experimental not working right now. Probably even delete test = Logger("<WATCHDOGWATCHDOG> ") def WatchdogReset(): test.notice("WATCHDOG has been reseted") with open("/dev/watchdog", "w") as f: f.write("1") #End of Experimental # Other subprocesses MiniHTTPServer().start() FileJanitor().start() FTPHandler().start() GraphDownloader().start() # Detection subprocesses img_queue = Queue(maxsize = 1) rst_queue = Queue(maxsize = conf.result_buffer) ImageFeed(img_queue).start() ResultHandler(rst_queue).start() while True: # CUDA / TensorFlow need to be in the main process PredictionPipeline(img_queue, rst_queue).predict() systemd.daemon.notify("WATCHDOG=1")
问题排查与解决方案
1. Systemd服务配置修正
- 将
Type=simple改为Type=notify:使用systemd.daemon.notify发送信号时,必须将服务类型设置为notify,否则systemd不会监听进程的通知信号。Type=simple模式下systemd仅检查进程是否存活,不处理通知,直接导致Watchdog机制失效。 - 移除
WatchdogTimestamp=30:该参数用于记录Watchdog触发时间,并非喂狗间隔,无需额外设置。 - 确保普通用户权限正常:默认情况下普通用户进程可正常使用
systemd.daemon.notify,若存在AppArmor/SELinux限制,需针对性调整,但优先修改服务类型测试。
修正后的服务配置:
[Unit] Description=Picturedetection Main application Wants=network-online.target After=network-online.target [Service] Type=notify User=user WorkingDirectory=/home/user/detection/ ExecStart=/usr/bin/python3 /home/user/detection/picturedetection.py Environment=TF_CUDNN_USE_AUTOTUNE=0 WatchdogSec=30 Restart=always [Install] WantedBy=multi-user.target
2. Python代码修正
- 明确导入systemd daemon模块:将
import systemd改为from systemd import daemon,避免隐式调用可能引发的潜在问题。 - 检查
predict()方法执行时长:如果PredictionPipeline.predict()执行时间超过30秒,循环内的喂狗操作会被延迟,触发Watchdog重启。需确保该方法在30秒内完成,或在方法内部添加喂狗逻辑(若为长期运行函数)。 - 移除硬件看门狗写入逻辑:写入
/dev/watchdog需要root权限,且与systemd软件Watchdog机制冲突,仅保留daemon.notify("WATCHDOG=1")即可。
修正后的关键代码片段:
# 替换导入语句 from systemd import daemon # ... 其他代码 ... while True: # CUDA / TensorFlow need to be in the main process # 若predict()执行时长可能超30秒,需在内部添加喂狗逻辑 PredictionPipeline(img_queue, rst_queue).predict() daemon.notify("WATCHDOG=1")
3. 验证步骤
- 重载并重启服务:
sudo systemctl daemon-reload && sudo systemctl restart picturedetection.service - 查看服务状态:
systemctl status picturedetection.service,确认无启动错误 - 实时查看日志:
journalctl -u picturedetection.service -f,观察是否有Watchdog通知记录或超时告警
疑问解答:多线程程序主进程卡住时,服务能否正常重启?
可以。Systemd Watchdog机制基于主进程状态判断:
- 若主进程陷入无限循环或阻塞在系统调用,无法按时发送
WATCHDOG=1通知,超过WatchdogSec时长后,systemd会判定服务异常并触发重启。 - 子线程/子进程的运行状态不影响systemd对主进程的判断,只要主进程无法完成喂狗操作,就会被强制重启。
内容的提问来源于stack exchange,提问作者ArtHax
相关产品推荐
相关产品推荐

