使用os.walk结合multiprocessing Pool时搜索目录错误问题排查
问题描述
我需要遍历二级驱动器(W:\)上的大型文件夹/文件结构,查找所有名称包含特定字符串的文件。使用os.walk可以实现,但耗时过长(超过20分钟)。为提升速度尝试使用Pool实现多进程,却发现返回的匹配结果来自本地C:\而非W:\。请问我哪里出错导致搜索了错误的目录?我知晓也可以使用多线程,但更希望聚焦于使用Pool的多进程解决方案。
def search_folders(root_dir): locations = {'File': [], 'Location': []} for subdir, dirs, files in os.walk(root_dir): for file in files: if "draft" in file: path = os.path.join(subdir, file) a = path.split("\\") file = a[len(a)-1] path = path.replace("/", "\\") locations['File'].append(file) locations['Location'].append(path) print(f"Doc found at: {path}") return locations if __name__ == '__main__': print('Started...') rootdir = 'W\\\\Inventory' found_items = {'File': [], 'Location': []} with Pool(14) as p: for subdir, dirs, files in os.walk(rootdir): for result in p.starmap(search_folders, subdir): found_items.update(result) p.close() p.join() df = pd.DataFrame.from_dict(found_items) df.to_excel('C:/temp/found.xlsx', index=False)
错误原因分析
- 路径写法错误:
'W\\\\Inventory'会被解析为W\\Inventory,多出来的反斜杠导致路径识别异常,应该写成'W:\\Inventory'或用原始字符串r'W:\Inventory'。 - 多进程调用逻辑完全错误:
- 主进程里提前执行
os.walk(rootdir),然后把遍历得到的subdir(子目录字符串)传给p.starmap。subdir是字符串,会被拆分成单个字符作为参数传递给search_folders,比如'W:\Inventory\Sub1'会拆成['W', ':', '\', 'I', ...]。 search_folders接收的root_dir变成单个字符,此时os.walk会默认从当前工作目录(通常是C盘路径)开始遍历,这就是结果来自C盘的原因。
- 主进程里提前执行
- 结果合并错误:用
update合并字典会直接替换列表,而不是追加内容,导致最终结果只保留最后一个进程的输出。
修正后的代码
import os from multiprocessing import Pool import pandas as pd def search_folders(root_dir): locations = {'File': [], 'Location': []} for subdir, _, files in os.walk(root_dir): for file in files: if "draft" in file: full_path = os.path.normpath(os.path.join(subdir, file)) locations['File'].append(file) locations['Location'].append(full_path) print(f"Doc found at: {full_path}") return locations if __name__ == '__main__': print('Started...') # 用原始字符串避免转义问题 root_dir = r'W:\Inventory' found_items = {'File': [], 'Location': []} # 获取根目录下的所有一级子目录,作为多进程的任务单元 task_dirs = [] for entry in os.scandir(root_dir): if entry.is_dir(): task_dirs.append(entry.path) # 如果根目录下没有子目录,直接将根目录作为任务 if not task_dirs: task_dirs.append(root_dir) # 多进程分配任务,每个进程处理一个子目录 with Pool(processes=14) as p: results = p.map(search_folders, task_dirs) # 合并所有进程的结果 for res in results: found_items['File'].extend(res['File']) found_items['Location'].extend(res['Location']) df = pd.DataFrame.from_dict(found_items) df.to_excel(r'C:\temp\found.xlsx', index=False) print("搜索完成,结果已导出")
关键修正点
- 路径规范:使用原始字符串
r'W:\Inventory'避免转义错误,确保正确指向W盘目录。 - 任务拆分:直接获取根目录下的一级子目录作为多进程任务,每个进程独立遍历一个子目录,避免主进程与子进程的路径混乱。
- 结果合并:用
extend代替update,实现列表内容的追加合并,保留所有搜索结果。 - 路径格式化:用
os.path.normpath统一路径格式,避免斜杠混用问题。
内容的提问来源于stack exchange,提问作者Paulg
相关产品推荐
相关产品推荐

