You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python并行化for循环下载文件遇到意外问题求助

问题原因

报错和核心数无关,是 multiprocessing 启动机制导致的:

  • Windows、macOS 平台下 multiprocessing 默认使用 spawn 方式启动子进程,子进程会重新导入整个父模块,没有入口保护的代码会被重复执行,所以出现打印重复的问题
  • 子进程导入模块时会再次触发创建进程的逻辑,在自身启动完成前就尝试开新进程,就会抛出你遇到的报错
  • 你传的 requests Response 对象 ret 本身也无法在 spawn 模式下被序列化传递给子进程,也是潜在报错点
适配模块调用的解决方案

不需要把并行逻辑放在当前模块的 __main__ 下,按以下步骤修改即可:

第一步:修改主调脚本添加入口保护

只需要调用模块的顶层脚本加保护即可,不影响模块的封装逻辑:

import this_code 
import another1_code 
import another2_code 

# 所有执行逻辑都放在 __main__ 保护块内
if __name__ == "__main__":
    #Step1
    another1_code.main()

    #Step2
    c_DateNow, c_DateIni, c_DateFin = another2_code.main()

    #Step3
    this_code.main(c_DateNow, c_DateIni, c_DateFin)

    #step4
    ## More code

第二步:修改下载模块用进程池实现并发

用 multiprocessing.Pool 自动管理进程生命周期,同时可自定义并发数,避免同时启动25个进程占用过多资源,还解决了对象序列化问题:

import os 
import requests
from multiprocessing import Pool, cpu_count

dataset="dataset_name"

# 下载函数修改为接收单参数拆包,只传可序列化的cookies
def down_file(task_args):
    dspath, file, savepath, cookies = task_args
    webfilename = dspath + file
    file_base = os.path.basename(file)
    local_file = os.path.join(savepath, file_base)
    print('...Downloading', file_base)
 
    req = requests.get(webfilename, cookies=cookies, allow_redirects=True, stream=True)
    filesize = int(req.headers['Content-length'])
    with open(local_file, 'wb') as outfile:
        chunk_size = 1048576
        for chunk in req.iter_content(chunk_size=chunk_size):
            outfile.write(chunk)
    return file_base

def download_files(filelist, c_DateNow, path_script, max_workers=None):
    # 自定义并发数,默认取CPU核心数,IO密集型任务最多设为核心数的2倍即可
    if max_workers is None:
        max_workers = min(cpu_count(), 8) # 限制最多同时8个下载,避免对站点造成压力
    ## Authenticate    
    url = 'url'
    values = {'email' : 'email', 'passwd' : "password", 'action' : 'login'}
    ret = requests.post(url, data=values)
    # 只提取需要的cookies传递,不要传整个Response对象
    cookies = ret.cookies

    ## Path to files
    dspath = 'datasetwebpath'
    savepath = os.path.join(path_script, dataset, c_DateNow)
    os.makedirs(savepath, exist_ok = True)

    # 构造任务参数列表
    task_list = [(dspath, file, savepath, cookies) for file in filelist]
    
    # 进程池执行任务
    with Pool(processes=max_workers) as pool:
        for finished_file in pool.imap_unordered(down_file, task_list):
            print(f"下载完成:{finished_file}")

def main(c_DateNow, c_DateIni, c_DateFin, path_script):    
    ## Other code
    files=["list of web file addresses"] 
    print("   ...Files being downladed\n     ", "\n      ".join(files), "\n")

    ## Download files
    download_files(files, c_DateNow, path_script)

可选Linux/macOS兼容方案

如果运行环境是Linux/macOS,可以在download_files函数开头加一行代码直接切换为fork启动方式,不需要修改主调脚本:

import multiprocessing
multiprocessing.set_start_method('fork', force=True)

fork方式直接复制父进程内存,不会重新导入模块,不会出现重复执行和启动报错的问题。

内容的提问来源于stack exchange,提问作者M.O.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 10:36:05