Python 3文件索引并行化改造需求:全机文件遍历提速
Got it, let's fix up your file indexing code to run in parallel—traversing an entire filesystem sequentially can be painfully slow, especially on large drives. Here's how to refactor it using Python's concurrent.futures module to leverage multiple CPU cores, plus some tweaks to avoid common pitfalls like race conditions when writing to the output file.
First, let's break down the bottlenecks in your original code:
- It uses a single thread/process to crawl every directory, wasting potential parallelism
- Directly writing to the file mid-traversal means all I/O operations are sequential too
- No error handling for restricted directories (which will crash the script on most systems)
Refactored Parallel Code
import os import time from concurrent.futures import ProcessPoolExecutor, as_completed def collect_file_paths(root_dir): """Recursively collect absolute paths of all files in a directory.""" file_paths = [] try: for dir_name, _, file_list in os.walk(root_dir): for fname in file_list: full_path = os.path.abspath(os.path.join(dir_name, fname)) file_paths.append(full_path.encode('utf-8')) # Match original encoding behavior except PermissionError: # Skip directories we don't have access to print(f"Permission denied: {root_dir}") except Exception as e: print(f"Error processing {root_dir}: {str(e)}") return file_paths if __name__ == "__main__": start = time.time() root_dir = '/' output_file = "test.txt" # Gather top-level directories to split parallel work top_level_dirs = [] try: with os.scandir(root_dir) as entries: for entry in entries: if entry.is_dir(follow_symlinks=False): top_level_dirs.append(entry.path) except PermissionError: print(f"Permission denied for root directory: {root_dir}") # Include the root directory itself in case it has files directly top_level_dirs.append(root_dir) # Use ProcessPoolExecutor (better than ThreadPool for filesystem tasks) all_file_paths = [] with ProcessPoolExecutor() as executor: # Submit all directory crawl tasks to the pool futures = {executor.submit(collect_file_paths, dir_path): dir_path for dir_path in top_level_dirs} # Collect results as they finish processing for future in as_completed(futures): dir_path = futures[future] try: paths = future.result() all_file_paths.extend(paths) except Exception as e: print(f"Failed to process {dir_path}: {str(e)}") # Write all collected paths to the file in one batch (avoids race conditions) with open(output_file, "wb") as handle: # Binary mode matches encoded paths handle.write(b'\n'.join(all_file_paths) + b'\n') end = time.time() print(f"Total time elapsed: {end - start:.2f} seconds")
Key Improvements Explained
- Parallel Crawling: We split the work by top-level directories, using
ProcessPoolExecutorto assign each directory to a separate process. This bypasses Python's GIL and lets multiple cores handle traversal simultaneously. - Safe File Writing: Instead of writing from each process, we collect all file paths in the main thread first, then write them all at once. This eliminates race conditions that would cause garbled output if multiple processes wrote to the file at the same time.
- Error Resilience: Added checks for
PermissionErrorand other exceptions, so the script doesn't crash when hitting restricted system directories. - Consistent Encoding: Matches your original code's
utf-8encoding and uses binary mode for writing to avoid encoding mismatches.
Why ProcessPool Instead of ThreadPool?
While ThreadPoolExecutor works for simple I/O-bound tasks, filesystem traversal involves a mix of I/O and CPU work (parsing directory structures, resolving paths). Using processes avoids GIL limitations and tends to deliver better performance on multi-core systems for this kind of task.
内容的提问来源于stack exchange,提问作者Chris

