Python 2.7基于scandir遍历子目录的实现优化与局限性问询
Hey there! I totally get why you're ditching os.walk here—2 million files is no joke, and that method's repeated system calls kill efficiency. Let's break down how to optimize this with scandir in Python 2.7, plus talk about common pitfalls and limitations.
scandir for Python 2.7 First, a quick note: since scandir isn't built into Python 2.7, you'll need to install a compatible version first (versions >=2.0 drop Python 2 support):
pip install scandir==1.10.0
The Efficient Implementation
This uses an iterative stack (instead of recursion) to avoid depth limits, leverages scandir's cached file attributes, and includes error handling for permission issues. For memory-sensitive scenarios, I'll also share a generator version.
List-Based Version (All Directories in Memory)
import scandir def get_all_directories(root_path): directories = [] # Use a stack to iterate instead of recursion (avoids RecursionError) stack = [root_path] while stack: current_dir = stack.pop() directories.append(current_dir) try: # Iterate with scandir: DirEntry objects have cached is_dir() for entry in scandir.scandir(current_dir): # Skip symlinks to directories to prevent loops if entry.is_dir(follow_symlinks=False): stack.append(entry.path) except OSError as e: print(f"Warning: Could not access {current_dir} - {e}") continue return directories
Generator Version (Memory-Friendly for Large Datasets)
If you're dealing with tens/hundreds of thousands of directories, a generator avoids loading everything into memory at once:
import scandir def get_all_directories_gen(root_path): stack = [root_path] while stack: current_dir = stack.pop() yield current_dir try: for entry in scandir.scandir(current_dir): if entry.is_dir(follow_symlinks=False): stack.append(entry.path) except OSError as e: print(f"Warning: Could not access {current_dir} - {e}") continue # Usage example: # for dir_path in get_all_directories_gen("/your/target/path"): # process(dir_path)
Why This Is Faster Than os.walk
- Cached Attributes:
scandirreturnsDirEntryobjects that cache file metadata (like whether it's a directory) from the initial system call.os.walkcallsos.listdirthenos.path.isdirfor every entry—this doubles the number of system calls, which is brutal with 2 million files. - Iterative Traversal: Recursion hits Python's default depth limit (~1000 levels), but a stack avoids this entirely.
Common Limitations of Basic Implementations (Including Your Existing Code)
If your current code has issues, these are the most likely culprits:
- Recursion Depth Errors: If you're using recursion instead of an iterative stack, deep directory trees will crash with
RecursionError. - Symbol Link Loops: Without
follow_symlinks=False, a symlink pointing to a parent directory will cause infinite traversal, eating up memory and CPU. - Missing Error Handling: Permission denied errors or broken symlinks will crash the entire traversal instead of being skipped.
- Memory Bloat: Storing all directories in a list can consume gigabytes of memory if you have hundreds of thousands of directories—use the generator version instead.
- Unnecessary File Checks: If your code still calls
os.path.isdiron entries fromscandir, you're wasting the cached attribute advantage.
Final Tips
- Always test with
follow_symlinks=Falseunless you explicitly need to traverse symlinked directories (and even then, add checks to avoid loops). - For extremely large directory trees, the generator version is non-negotiable—it keeps memory usage low regardless of how many directories you have.
内容的提问来源于stack exchange,提问作者Jens Hiestermann

