You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多进程环境下无法向列表存储/追加值的问题求助

Python多进程中无法向列表追加值的问题解决

问题根源

多进程模式下,每个子进程会独立复制一份主进程的内存数据。你在子进程里对failed_files的修改,只作用于子进程自己的副本,主进程中的原列表完全不受影响,所以最后打印的始终是空列表。另外,就算你取消# Global failed_files的注释声明全局变量也没用——global仅能作用于同一进程内的变量,跨进程不生效。


修正方案1:使用可共享内存的列表

通过multiprocessing.Manager创建支持跨进程共享的列表,所有子进程对该列表的修改会同步到主进程:

import multiprocessing

# 创建Manager管理的共享列表
manager = multiprocessing.Manager()
failed_files = manager.list()

def get_text_pdfs(filename):
    # 省略原有业务逻辑
    coll = []
    if len(coll) == 15:
        insert_into_db()
    else:
        failed_files.append(filename)

list_of_pdfs = get_pdfs_list(path)
if list_of_pdfs is not None:
    # 用with语句自动管理进程池的生命周期
    with multiprocessing.Pool(5) as p:
        p.map(get_text_pdfs, list_of_pdfs)
print(failed_files)

修正方案2:通过返回值收集结果

不依赖共享内存,让子进程返回处理状态,主进程统一收集失败的文件名,这种方式更安全简洁:

import multiprocessing

def get_text_pdfs(filename):
    # 省略原有业务逻辑
    coll = []
    if len(coll) == 15:
        insert_into_db()
        return None  # 处理成功返回None
    else:
        return filename  # 处理失败返回文件名

list_of_pdfs = get_pdfs_list(path)
if list_of_pdfs is not None:
    with multiprocessing.Pool(5) as p:
        results = p.map(get_text_pdfs, list_of_pdfs)
    # 过滤掉成功项,收集失败文件
    failed_files = [f for f in results if f is not None]
print(failed_files)

额外代码规范建议

  • 条件判断if(list_of_pdfs!=None)建议改为if list_of_pdfs is not None,这是Python的标准写法。
  • 使用Pool时优先用with语句,无需手动调用close()和join(),避免资源泄漏。

内容的提问来源于stack exchange,提问作者Vignesh Vangala

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 00:25:21