You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python线程挂起时如何抛出异常并自动重启脚本?

解决concurrent.futures线程挂起超时后自动重启脚本的问题

问题背景

我有一个全天候运行的Python脚本,用concurrent.futures.ThreadPoolExecutor处理请求,但部分线程发起请求后无响应,导致脚本卡住。用hanging-threads 2.0.5可以打印挂起线程的调用栈:

Thread 139646566659840 "ThreadPoolExecutor-666849_1" hangs -
File "/usr/lib/python3.9/threading.py", line 912, in _bootstrap
self._bootstrap_inner()
File "/usr/lib/python3.9/threading.py", line 954, in _bootstrap_inner
self.run()
File "/usr/lib/python3.9/threading.py", line 892, in run
self._target(*self._args, **self._kwargs)
File "/usr/lib/python3.9/concurrent/futures/thread.py", line 77, in _worker
work_item.run()
File "/usr/lib/python3.9/concurrent/futures/thread.py", line 52, in run
result = self.fn(*self.args, **self.kwargs)

需求是在线程超时无响应时触发脚本自动重启,但直接给future.result()设timeout无法取消已运行的线程——因为Python线程无法被强制终止。


解决方案

方案1:用进程池替代线程池(推荐)

进程拥有独立内存空间,支持强制终止,改用ProcessPoolExecutor可解决线程无法取消的问题:

from concurrent.futures import ProcessPoolExecutor
import sys

def task():
    # 你的请求处理逻辑(可能卡住的代码)
    ...

def main():
    with ProcessPoolExecutor(max_workers=4) as executor:
        while True:
            future = executor.submit(task)
            try:
                # 设置任务超时时间,例如30秒
                result = future.result(timeout=30)
                # 处理正常返回结果
                ...
            except TimeoutError:
                # 终止超时的子进程
                future.cancel()
                print("任务超时,终止进程并重启脚本")
                # 重启当前脚本
                sys.execv(sys.executable, ['python'] + sys.argv)

if __name__ == "__main__":
    main()

触发TimeoutError时,future.cancel()会向子进程发送终止信号,真正停止任务,随后调用sys.execv完成脚本重启。

方案2:基于hanging-threads监控触发重启

如果不想改动线程池的使用方式,可利用hanging-threads的监控能力,检测到线程挂起时主动重启脚本:

from hanging_threads import start_monitoring
import sys
import time
from concurrent.futures import ThreadPoolExecutor

def on_hang(thread_id, thread_name, call_stack):
    print(f"线程 {thread_id} ({thread_name}) 挂起,调用栈:\n{call_stack}")
    print("触发脚本重启")
    # 重启当前脚本
    sys.execv(sys.executable, ['python'] + sys.argv)

# 启动监控:线程在同一栈帧停留超过30秒判定为挂起,每10秒检测一次
start_monitoring(seconds_frozen=30, check_interval=10, on_freeze=on_hang)

def task():
    # 你的请求处理逻辑
    ...

def main():
    with ThreadPoolExecutor(max_workers=4) as executor:
        while True:
            executor.submit(task)
            time.sleep(1)

if __name__ == "__main__":
    main()

这种方式通过后台线程定时检测所有线程的栈帧,一旦发现挂起线程就触发重启逻辑,无需修改原有任务提交流程。


说明

Python线程受GIL限制,无法强制终止运行中的线程,因此要么用进程池实现可终止的任务,要么通过外部监控触发脚本整体重启,两种方案都能解决脚本卡住的问题。

内容的提问来源于stack exchange,提问作者PyNoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 13:55:23