You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用os.walk结合multiprocessing Pool时搜索目录错误问题排查

问题描述

我需要遍历二级驱动器(W:\)上的大型文件夹/文件结构,查找所有名称包含特定字符串的文件。使用os.walk可以实现,但耗时过长(超过20分钟)。为提升速度尝试使用Pool实现多进程,却发现返回的匹配结果来自本地C:\而非W:\。请问我哪里出错导致搜索了错误的目录?我知晓也可以使用多线程,但更希望聚焦于使用Pool的多进程解决方案。

def search_folders(root_dir):

    locations = {'File': [],
                 'Location': []}

    for subdir, dirs, files in os.walk(root_dir):

        for file in files:
            if "draft" in file:
                path = os.path.join(subdir, file)
                a = path.split("\\")
                file = a[len(a)-1]
                path = path.replace("/", "\\")
                locations['File'].append(file)
                locations['Location'].append(path)
                print(f"Doc found at: {path}")
    return locations


if __name__ == '__main__':
    print('Started...')
    rootdir = 'W\\\\Inventory'

    found_items = {'File': [],
                   'Location': []}

    with Pool(14) as p:
        for subdir, dirs, files in os.walk(rootdir):
            for result in p.starmap(search_folders, subdir):
                found_items.update(result)

    p.close()
    p.join()

    df = pd.DataFrame.from_dict(found_items)
    df.to_excel('C:/temp/found.xlsx', index=False)
错误原因分析
  1. 路径写法错误:'W\\\\Inventory'会被解析为W\\Inventory,多出来的反斜杠导致路径识别异常,应该写成'W:\\Inventory'或用原始字符串r'W:\Inventory'。
  2. 多进程调用逻辑完全错误:
    • 主进程里提前执行os.walk(rootdir),然后把遍历得到的subdir(子目录字符串)传给p.starmap。subdir是字符串,会被拆分成单个字符作为参数传递给search_folders,比如'W:\Inventory\Sub1'会拆成['W', ':', '\', 'I', ...]。
    • search_folders接收的root_dir变成单个字符,此时os.walk会默认从当前工作目录(通常是C盘路径)开始遍历,这就是结果来自C盘的原因。
  3. 结果合并错误:用update合并字典会直接替换列表,而不是追加内容,导致最终结果只保留最后一个进程的输出。
修正后的代码
import os
from multiprocessing import Pool
import pandas as pd

def search_folders(root_dir):
    locations = {'File': [], 'Location': []}
    for subdir, _, files in os.walk(root_dir):
        for file in files:
            if "draft" in file:
                full_path = os.path.normpath(os.path.join(subdir, file))
                locations['File'].append(file)
                locations['Location'].append(full_path)
                print(f"Doc found at: {full_path}")
    return locations

if __name__ == '__main__':
    print('Started...')
    # 用原始字符串避免转义问题
    root_dir = r'W:\Inventory'
    
    found_items = {'File': [], 'Location': []}
    
    # 获取根目录下的所有一级子目录,作为多进程的任务单元
    task_dirs = []
    for entry in os.scandir(root_dir):
        if entry.is_dir():
            task_dirs.append(entry.path)
    # 如果根目录下没有子目录,直接将根目录作为任务
    if not task_dirs:
        task_dirs.append(root_dir)
    
    # 多进程分配任务,每个进程处理一个子目录
    with Pool(processes=14) as p:
        results = p.map(search_folders, task_dirs)
    
    # 合并所有进程的结果
    for res in results:
        found_items['File'].extend(res['File'])
        found_items['Location'].extend(res['Location'])
    
    df = pd.DataFrame.from_dict(found_items)
    df.to_excel(r'C:\temp\found.xlsx', index=False)
    print("搜索完成,结果已导出")
关键修正点
  • 路径规范:使用原始字符串r'W:\Inventory'避免转义错误,确保正确指向W盘目录。
  • 任务拆分:直接获取根目录下的一级子目录作为多进程任务,每个进程独立遍历一个子目录,避免主进程与子进程的路径混乱。
  • 结果合并:用extend代替update,实现列表内容的追加合并,保留所有搜索结果。
  • 路径格式化:用os.path.normpath统一路径格式,避免斜杠混用问题。

内容的提问来源于stack exchange,提问作者Paulg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 08:35:18