Python并行化for循环下载文件遇到意外问题求助
问题原因
报错和核心数无关,是 multiprocessing 启动机制导致的:
- Windows、macOS 平台下 multiprocessing 默认使用
spawn方式启动子进程,子进程会重新导入整个父模块,没有入口保护的代码会被重复执行,所以出现打印重复的问题 - 子进程导入模块时会再次触发创建进程的逻辑,在自身启动完成前就尝试开新进程,就会抛出你遇到的报错
- 你传的 requests Response 对象
ret本身也无法在spawn模式下被序列化传递给子进程,也是潜在报错点
适配模块调用的解决方案
不需要把并行逻辑放在当前模块的 __main__ 下,按以下步骤修改即可:
第一步:修改主调脚本添加入口保护
只需要调用模块的顶层脚本加保护即可,不影响模块的封装逻辑:
import this_code import another1_code import another2_code # 所有执行逻辑都放在 __main__ 保护块内 if __name__ == "__main__": #Step1 another1_code.main() #Step2 c_DateNow, c_DateIni, c_DateFin = another2_code.main() #Step3 this_code.main(c_DateNow, c_DateIni, c_DateFin) #step4 ## More code
第二步:修改下载模块用进程池实现并发
用 multiprocessing.Pool 自动管理进程生命周期,同时可自定义并发数,避免同时启动25个进程占用过多资源,还解决了对象序列化问题:
import os import requests from multiprocessing import Pool, cpu_count dataset="dataset_name" # 下载函数修改为接收单参数拆包,只传可序列化的cookies def down_file(task_args): dspath, file, savepath, cookies = task_args webfilename = dspath + file file_base = os.path.basename(file) local_file = os.path.join(savepath, file_base) print('...Downloading', file_base) req = requests.get(webfilename, cookies=cookies, allow_redirects=True, stream=True) filesize = int(req.headers['Content-length']) with open(local_file, 'wb') as outfile: chunk_size = 1048576 for chunk in req.iter_content(chunk_size=chunk_size): outfile.write(chunk) return file_base def download_files(filelist, c_DateNow, path_script, max_workers=None): # 自定义并发数,默认取CPU核心数,IO密集型任务最多设为核心数的2倍即可 if max_workers is None: max_workers = min(cpu_count(), 8) # 限制最多同时8个下载,避免对站点造成压力 ## Authenticate url = 'url' values = {'email' : 'email', 'passwd' : "password", 'action' : 'login'} ret = requests.post(url, data=values) # 只提取需要的cookies传递,不要传整个Response对象 cookies = ret.cookies ## Path to files dspath = 'datasetwebpath' savepath = os.path.join(path_script, dataset, c_DateNow) os.makedirs(savepath, exist_ok = True) # 构造任务参数列表 task_list = [(dspath, file, savepath, cookies) for file in filelist] # 进程池执行任务 with Pool(processes=max_workers) as pool: for finished_file in pool.imap_unordered(down_file, task_list): print(f"下载完成:{finished_file}") def main(c_DateNow, c_DateIni, c_DateFin, path_script): ## Other code files=["list of web file addresses"] print(" ...Files being downladed\n ", "\n ".join(files), "\n") ## Download files download_files(files, c_DateNow, path_script)
可选Linux/macOS兼容方案
如果运行环境是Linux/macOS,可以在download_files函数开头加一行代码直接切换为fork启动方式,不需要修改主调脚本:
import multiprocessing multiprocessing.set_start_method('fork', force=True)
fork方式直接复制父进程内存,不会重新导入模块,不会出现重复执行和启动报错的问题。
内容的提问来源于stack exchange,提问作者M.O.
相关产品推荐
相关产品推荐

