You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速Python中os.walk()结合glob的网络驱动器文件查找代码?

Great question—dealing with slow network drives and massive file sets is always tricky, especially when you only need a tiny fraction of the files. Let's first unpack why your current code is slow: os.walk iterates through every single directory (even ones that can't possibly contain your target files), and each call to glob.glob adds extra network IO overhead as it scans the current directory. When dealing with 6TB of data and hundreds of thousands of files, all that sequential network waiting adds up to hours of runtime.

Here are several proven optimizations to speed up your file search:

1. Simplify with Recursive Glob or Pathlib's rglob

Python 3.5+ added recursive support to glob.glob, and Python 3.4+ introduced pathlib which has a cleaner rglob method. Both can replace your nested loop with a single call, and in some cases, they're optimized to reduce redundant IO operations.

Using glob.glob recursively:

import glob
import os

pattern = "*20210929.hdf5"
files = glob.glob(os.path.join(directory, "**", pattern), recursive=True)

Using pathlib.Path.rglob (more readable):

from pathlib import Path

pattern = "*20210929.hdf5"
files = list(Path(directory).rglob(pattern))
# Convert to strings if needed: files = [str(f) for f in files]

2. Parallelize Scans to Leverage Your Xeon CPU

Your high-performance CPU is likely idle while waiting for network IO. Threading is perfect here (since this is an IO-bound task) — you can split the directory tree into chunks and scan multiple subdirectories at the same time.

Here's how to implement this with concurrent.futures.ThreadPoolExecutor:

import os
import glob
from concurrent.futures import ThreadPoolExecutor

def scan_directory(subdir, pattern):
    return glob.glob(os.path.join(subdir, "**", pattern), recursive=True)

pattern = "*20210929.hdf5"
top_level_dirs = [os.path.join(directory, d) for d in os.listdir(directory) if os.path.isdir(os.path.join(directory, d))]

# Adjust max_workers based on your network capacity (start with 4-8, test what works)
with ThreadPoolExecutor(max_workers=8) as executor:
    results = executor.map(scan_directory, top_level_dirs, [pattern]*len(top_level_dirs))

# Flatten the list of results
files = [item for sublist in results for item in sublist]

This way, you're not waiting for one directory scan to finish before starting the next — multiple scans run in parallel, cutting down total waiting time.

3. Skip Unlikely Directories Early

If you can identify directories that can't contain your target files, skip them entirely to avoid unnecessary network IO. For example:

  • Skip hidden directories (names starting with . or _)
  • Skip directories whose names don't relate to your date pattern (e.g., if your file is dated 20210929, skip directories named 2020* or 202110*)

Modify your original loop (or the parallel version) to add a check:

import os
import glob

pattern = "*20210929.hdf5"
target_date = "20210929"
files = []

for dirpath, _, _ in os.walk(directory):
    # Skip hidden directories
    if os.path.basename(dirpath).startswith((".", "_")):
        continue
    # Skip directories that don't contain the target date (if your directory structure uses dates)
    if target_date not in dirpath:
        continue
    files.extend(glob.glob(os.path.join(dirpath, pattern)))

This can drastically reduce the number of directories you need to scan.

4. Use System-Native Search Tools (Fastest Option)

Operating systems have optimized, low-level tools for file search that often outperform Python's built-ins, especially on network drives. These tools are written in compiled languages and leverage OS-level caching and optimizations.

For Windows:

Use the dir command with recursive search:

import subprocess
import os

pattern = "*20210929.hdf5"
command = f'dir /s /b "{os.path.join(directory, pattern)}"'
result = subprocess.run(command, shell=True, capture_output=True, text=True)
files = [line.strip() for line in result.stdout.splitlines() if line.strip()]

For Linux/macOS:

Use the find command:

import subprocess
import os

pattern = "*20210929.hdf5"
command = f'find "{directory}" -type f -name "{pattern}"'
result = subprocess.run(command, shell=True, capture_output=True, text=True)
files = [line.strip() for line in result.stdout.splitlines() if line.strip()]

This is often the fastest method because it bypasses Python's higher-level file system abstractions and uses the OS's native, optimized search logic.

Final Recommendations

  • If you need cross-platform compatibility, start with the parallelized rglob/glob approach.
  • If you're on a single OS, the system-native tool method will likely give you the biggest speed boost.
  • Always test with smaller subsets first to tune parameters like thread count or directory filters.

内容的提问来源于stack exchange,提问作者jB777

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 14:52:29