You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python pathlib glob高效并行搜索大型目录结构?

优化超大型目录文件搜索的几种方案

1. 多线程结合os.scandir并行遍历

os.scandir是Python 3.5+引入的高效目录遍历接口,比pathlib.rglob更快——它直接返回带缓存信息的目录项,大幅减少系统调用次数。结合多线程(IO密集场景适配性更好)拆分遍历任务,能显著提升速度:

import os
from concurrent.futures import ThreadPoolExecutor

def search_in_dir(dir_path, target_file):
    matches = []
    for entry in os.scandir(dir_path):
        if entry.is_file(follow_symlinks=False) and entry.name == target_file:
            matches.append(entry.path)
        elif entry.is_dir(follow_symlinks=False):
            matches.extend(search_in_dir(entry.path, target_file))
    return matches

def parallel_search(base_dir, target_file, max_workers=8):
    # 收集一级子目录和根目录作为并行任务入口
    initial_dirs = [base_dir]
    for entry in os.scandir(base_dir):
        if entry.is_dir(follow_symlinks=False):
            initial_dirs.append(entry.path)
    
    matches = []
    with ThreadPoolExecutor(max_workers=max_workers) as executor:
        futures = [executor.submit(search_in_dir, d, target_file) for d in initial_dirs]
        for future in futures:
            matches.extend(future.result())
    return matches

# 使用示例
if __name__ == "__main__":
    base_dir = "/path/to/large/directory"
    target_file = "ind_stat.zpkl"
    files = parallel_search(base_dir, target_file)

2. 多进程拆分任务(适配超大规模目录)

如果目录结构极端庞大,多进程可利用多核优势抵消进程切换开销,进一步提升遍历效率:

import os
from multiprocessing import Pool

def search_in_dir(dir_path, target_file):
    matches = []
    for entry in os.scandir(dir_path):
        if entry.is_file(follow_symlinks=False) and entry.name == target_file:
            matches.append(entry.path)
        elif entry.is_dir(follow_symlinks=False):
            matches.extend(search_in_dir(entry.path, target_file))
    return matches

def multiprocess_search(base_dir, target_file, processes=4):
    initial_dirs = [base_dir]
    for entry in os.scandir(base_dir):
        if entry.is_dir(follow_symlinks=False):
            initial_dirs.append(entry.path)
    
    with Pool(processes=processes) as pool:
        results = pool.starmap(search_in_dir, [(d, target_file) for d in initial_dirs])
    
    matches = []
    for res in results:
        matches.extend(res)
    return matches

# 使用示例
if __name__ == "__main__":
    base_dir = "/path/to/large/directory"
    target_file = "ind_stat.zpkl"
    files = multiprocess_search(base_dir, target_file)

3. 用高效第三方库替代

  • glob2: 对遍历逻辑做了底层优化,比标准库glob更快,支持递归搜索:
    import glob2
    files = glob2.glob("/path/to/large/directory/**/ind_stat.zpkl", recursive=True)
    
  • fastglob: 基于C扩展实现,纯Python方案的速度天花板,适合极致性能需求:
    from fastglob import fastglob
    files = fastglob("/path/to/large/directory/**/ind_stat.zpkl", recursive=True)
    

4. 预缓存目录结构(适合重复搜索场景)

如果需要多次搜索同一目录,可提前缓存所有文件路径,后续直接查询缓存避免重复遍历:

# 首次运行生成缓存
import os
import json

def cache_dir_structure(base_dir, cache_path="dir_cache.json"):
    file_paths = []
    for root, _, files in os.walk(base_dir):
        for f in files:
            file_paths.append(os.path.join(root, f))
    with open(cache_path, "w") as f:
        json.dump(file_paths, f)

# 后续搜索直接读缓存
def search_from_cache(target_file, cache_path="dir_cache.json"):
    with open(cache_path, "r") as f:
        all_files = json.load(f)
    return [p for p in all_files if os.path.basename(p) == target_file]

注意事项

  • 多线程/多进程的并发数不要设置过大,否则会导致系统IO过载,建议线程数设为CPU核心数的2倍,进程数等于CPU核心数。
  • 关闭符号链接遍历(follow_symlinks=False),避免循环遍历和无效开销。
  • 若目标是网络文件系统(NFS/SMB),并行方案的优势会被网络IO瓶颈抵消,建议改用单线程高效遍历。

内容的提问来源于stack exchange,提问作者Roy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 12:05:07