You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何释放multiprocessing中res_list的内存?apply_async使用疑问

解决multiprocessing apply_async内存无法释放的问题

问题根源

你的代码存在两个核心问题导致内存无法有效释放:

  1. 双列表内存占用叠加:先把10000个AsyncResult对象全存入res_list,再统一调用get()将所有矩阵存入result列表——此时内存同时承载res_list和result,后者(矩阵集合)的内存占用远大于前者,所以即使res_list被设为None,内存也不会明显下降。
  2. AsyncResult的缓存残留:调用get()后,AsyncResult对象可能会缓存结果,直到被彻底回收。你在所有get()完成后才清理res_list,期间这些对象一直占用内存。

另外,你的代码缺少if __name__ == "__main__":保护,在Windows环境下会触发多进程启动错误,这是multiprocessing的标准使用要求。

解决方案

方案1:分批提交+分批清理内存

避免一次性保存所有AsyncResult和结果,分批处理后立即清理当前批次的资源:

from multiprocessing import Pool
import gc

def calculate(i, j):
    # 替换成你的矩阵生成逻辑
    return [[i+j for _ in range(100)] for _ in range(100)]

if __name__ == "__main__":
    pool = Pool(10)
    batch_size = 100  # 每次处理100个任务,可根据内存调整
    total_tasks = 10000
    result = []
    
    for batch_start in range(0, total_tasks, batch_size):
        res_batch = []
        # 提交当前批次任务
        for idx in range(batch_start, min(batch_start + batch_size, total_tasks)):
            i = idx // 100
            j = idx % 100
            res_batch.append(pool.apply_async(calculate, (i, j)))
        
        # 获取当前批次结果
        for res in res_batch:
            result.append(res.get())
        
        # 清理当前批次的AsyncResult,触发垃圾回收
        res_batch = None
        gc.collect()
    
    pool.close()
    pool.join()
    
    # 后续处理result...

方案2:用回调函数直接处理结果(无需保存所有结果)

如果不需要把所有矩阵存在内存里,可以用callback参数在结果返回时直接处理(比如写入文件/数据库),完全不用保存AsyncResult或result列表:

from multiprocessing import Pool

def calculate(i, j):
    return [[i+j for _ in range(100)] for _ in range(100)]

def handle_result(result):
    # 替换成你的结果处理逻辑,比如写入文件
    with open("matrix_results.txt", "a") as f:
        f.write(f"Matrix shape: {len(result)}x{len(result[0])}\n")

if __name__ == "__main__":
    pool = Pool(10)
    
    for i in range(100):
        for j in range(100):
            pool.apply_async(calculate, (i, j), callback=handle_result)
    
    pool.close()
    pool.join()

方案3:逐个清理AsyncResult(最小改动原代码)

在调用get()后立即删除对应的AsyncResult对象,减少内存占用:

from multiprocessing import Pool
import gc

def calculate(i, j):
    return [[i+j for _ in range(100)] for _ in range(100)]

if __name__ == "__main__":
    res_list = []
    pool = Pool(10)
    for i in range(100):
        for j in range(100):
            res_list.append(pool.apply_async(calculate, (i, j)))

    pool.close()
    pool.join()

    result = []
    # 逐个取出AsyncResult,get后删除,避免同时持有所有对象
    while res_list:
        # pop()从末尾取比pop(0)效率高
        res = res_list.pop()
        result.append(res.get())
        gc.collect()
    
    # 此时res_list已为空,无需额外赋值None

关键注意事项

  • 必须把多进程代码放在if __name__ == "__main__":块中,否则Windows系统会出现进程启动错误。
  • Python的垃圾回收是自动的,gc.collect()只是触发一次回收,不是强制立即释放所有内存,系统层面的内存回收可能有延迟。
  • 如果矩阵本身极大,即使分批处理也会占用大量内存,建议直接在回调里持久化结果,不要存在内存中。

内容的提问来源于stack exchange,提问作者xuu xuu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 22:45:33