You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Docker容器中multiprocessing.Queue.put()致进程崩溃问题求助

问题:Python多进程Queue在Linux/Docker中调用put()时进程意外终止

我在Docker容器中运行一个基于Python multiprocessing模块的测试程序,采用生产者-消费者模型,用multiprocessing.Queue实现进程间通信。主进程启动生产者进程后,后者调用Queue.put()时会被终止,且无任何异常抛出。该程序在本地macOS环境运行正常,但在Python 3.9.16的Docker基础镜像及常规Debian虚拟机中均出现此问题。


复现代码

Python测试代码

import multiprocessing as mp
import time
import traceback
from typing import Any, Optional

import psutil

base_file = "logs.txt"


def main() -> None:
    queue: Any = mp.Queue()
    print("Queue created")
    print("Starting producer process")

    p = mp.get_context("spawn").Process(target=producer, args=(queue,), daemon=True)
    p.start()
    print(f"Main: producer started: {p.pid}")

    alive = True
    while alive:
        alive = p.is_alive()
        print(f"Ha ha ha staying alive, producer: {p.is_alive()}")
        time.sleep(1)

    print("Every process is dead :(")

def producer(q: mp.Queue) -> None:
    with open(f"producer.{base_file}", "w") as f:
        print("Producer: started", file=f, flush=True)
        current_value: int = 0
        while True:
            print(f"Producer: Adding value {current_value} to queue", file=f, flush=True)
            try:
                q.put(current_value, block=False)
            except BaseException as e:
                print(f"Producer: exception: {e}", file=f, flush=True)
                print(f"{traceback.format_exc()}", file=f, flush=True)
                raise e
            print(f"Producer: Value {current_value} added to queue", file=f, flush=True)
            print("Producer: Sleeping for 1 second", file=f, flush=True)
            time.sleep(1)
            current_value += 1


if __name__ == "__main__":
    main()

Dockerfile

FROM python:3.9.16

RUN apt-get update && apt-get install -y gettext git mime-support && apt-get clean

RUN python3 -m pip install psutil

COPY ./multiprocessing_e2e.py /src/multiprocessing_e2e.py

WORKDIR /src

CMD ["python", "-u", "multiprocessing_e2e.py"]

问题原因与解决方法

核心原因

Linux环境下,multiprocessing.Queue依赖管道和内部线程实现跨进程通信。当主进程仅检查子进程存活状态但完全不消费队列数据时,队列的缓冲区会快速被填满。虽然代码中用block=False调用put(),理论上应抛出Queue.Full异常,但在spawn启动模式的Linux系统中,队列底层管道写端因缓冲区耗尽触发的内部逻辑会直接终止子进程,且这种终止不会被Python的BaseException捕获。

macOS的进程调度、队列实现逻辑与Linux存在差异,因此本地运行无此问题。

解决方法

  1. 主进程添加队列消费逻辑(根本解决):
    修改主进程代码,定期从队列中取数据,避免缓冲区被填满:

    def main() -> None:
        queue: Any = mp.Queue()
        print("Queue created")
        print("Starting producer process")
    
        p = mp.get_context("spawn").Process(target=producer, args=(queue,), daemon=True)
        p.start()
        print(f"Main: producer started: {p.pid}")
    
        alive = True
        while alive:
            alive = p.is_alive()
            # 新增消费逻辑
            try:
                item = queue.get(block=False)
                print(f"Main: Received {item} from queue")
            except mp.Queue.Empty:
                pass
            print(f"Ha ha ha staying alive, producer: {p.is_alive()}")
            time.sleep(1)
    
        print("Every process is dead :(")
    
  2. 移除block=False参数:
    如果不需要非阻塞写入,去掉该参数,让put()在队列满时阻塞等待,避免触发底层异常终止逻辑。

  3. 增大队列初始容量:
    创建Queue时指定maxsize参数(如mp.Queue(maxsize=1000)),临时缓解缓冲区耗尽问题,但无法替代消费逻辑。


内容的提问来源于stack exchange,提问作者brvh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 00:02:07