如何加速Python中os.walk()结合glob的网络驱动器文件查找代码?
Great question—dealing with slow network drives and massive file sets is always tricky, especially when you only need a tiny fraction of the files. Let's first unpack why your current code is slow: os.walk iterates through every single directory (even ones that can't possibly contain your target files), and each call to glob.glob adds extra network IO overhead as it scans the current directory. When dealing with 6TB of data and hundreds of thousands of files, all that sequential network waiting adds up to hours of runtime.
Here are several proven optimizations to speed up your file search:
1. Simplify with Recursive Glob or Pathlib's rglob
Python 3.5+ added recursive support to glob.glob, and Python 3.4+ introduced pathlib which has a cleaner rglob method. Both can replace your nested loop with a single call, and in some cases, they're optimized to reduce redundant IO operations.
Using glob.glob recursively:
import glob import os pattern = "*20210929.hdf5" files = glob.glob(os.path.join(directory, "**", pattern), recursive=True)
Using pathlib.Path.rglob (more readable):
from pathlib import Path pattern = "*20210929.hdf5" files = list(Path(directory).rglob(pattern)) # Convert to strings if needed: files = [str(f) for f in files]
2. Parallelize Scans to Leverage Your Xeon CPU
Your high-performance CPU is likely idle while waiting for network IO. Threading is perfect here (since this is an IO-bound task) — you can split the directory tree into chunks and scan multiple subdirectories at the same time.
Here's how to implement this with concurrent.futures.ThreadPoolExecutor:
import os import glob from concurrent.futures import ThreadPoolExecutor def scan_directory(subdir, pattern): return glob.glob(os.path.join(subdir, "**", pattern), recursive=True) pattern = "*20210929.hdf5" top_level_dirs = [os.path.join(directory, d) for d in os.listdir(directory) if os.path.isdir(os.path.join(directory, d))] # Adjust max_workers based on your network capacity (start with 4-8, test what works) with ThreadPoolExecutor(max_workers=8) as executor: results = executor.map(scan_directory, top_level_dirs, [pattern]*len(top_level_dirs)) # Flatten the list of results files = [item for sublist in results for item in sublist]
This way, you're not waiting for one directory scan to finish before starting the next — multiple scans run in parallel, cutting down total waiting time.
3. Skip Unlikely Directories Early
If you can identify directories that can't contain your target files, skip them entirely to avoid unnecessary network IO. For example:
- Skip hidden directories (names starting with
.or_) - Skip directories whose names don't relate to your date pattern (e.g., if your file is dated 20210929, skip directories named
2020*or202110*)
Modify your original loop (or the parallel version) to add a check:
import os import glob pattern = "*20210929.hdf5" target_date = "20210929" files = [] for dirpath, _, _ in os.walk(directory): # Skip hidden directories if os.path.basename(dirpath).startswith((".", "_")): continue # Skip directories that don't contain the target date (if your directory structure uses dates) if target_date not in dirpath: continue files.extend(glob.glob(os.path.join(dirpath, pattern)))
This can drastically reduce the number of directories you need to scan.
4. Use System-Native Search Tools (Fastest Option)
Operating systems have optimized, low-level tools for file search that often outperform Python's built-ins, especially on network drives. These tools are written in compiled languages and leverage OS-level caching and optimizations.
For Windows:
Use the dir command with recursive search:
import subprocess import os pattern = "*20210929.hdf5" command = f'dir /s /b "{os.path.join(directory, pattern)}"' result = subprocess.run(command, shell=True, capture_output=True, text=True) files = [line.strip() for line in result.stdout.splitlines() if line.strip()]
For Linux/macOS:
Use the find command:
import subprocess import os pattern = "*20210929.hdf5" command = f'find "{directory}" -type f -name "{pattern}"' result = subprocess.run(command, shell=True, capture_output=True, text=True) files = [line.strip() for line in result.stdout.splitlines() if line.strip()]
This is often the fastest method because it bypasses Python's higher-level file system abstractions and uses the OS's native, optimized search logic.
Final Recommendations
- If you need cross-platform compatibility, start with the parallelized
rglob/globapproach. - If you're on a single OS, the system-native tool method will likely give you the biggest speed boost.
- Always test with smaller subsets first to tune parameters like thread count or directory filters.
内容的提问来源于stack exchange,提问作者jB777

