Python中如何无需遍历全目录快速筛选指定日期之后的文件?
Great question—full directory traversals can be a total slog with large folders, right? The good news is we have options in Python that match (or come close to) the efficiency of find . -newermt on Unix-like systems. Let’s break down the best approaches:
1. Use os.scandir() (Python 3.5+) – Fast Native Python Approach
os.scandir() is a game-changer for directory traversal because it caches file metadata (like modification time) when it reads the directory, avoiding the extra os.stat() calls that slow down older methods like os.walk(). It still traverses the directory tree, but it’s the fastest pure-Python method available.
Here’s a recursive implementation that filters files newer than your target date:
import os from datetime import datetime def find_files_newer_than(root_dir, target_date_str): # Convert target date string to a Unix timestamp (seconds since epoch) target_timestamp = datetime.strptime(target_date_str, '%Y-%m-%d %H:%M:%S').timestamp() for entry in os.scandir(root_dir): if entry.is_file(follow_symlinks=False): # Access cached mtime directly from the DirEntry object if entry.stat().st_mtime > target_timestamp: yield entry.path elif entry.is_dir(follow_symlinks=False): # Recursively check subdirectories yield from find_files_newer_than(entry.path, target_date_str) # Example usage for file_path in find_files_newer_than('.', '2018-01-17 03:28:46'): print(file_path)
Why this works:
os.scandir()returnsDirEntryobjects that already hold stat info, so we don’t waste time re-querying the filesystem for each file.- It’s cross-platform (works on Windows, macOS, Linux) and doesn’t rely on external tools.
2. Call the System’s find Command (Unix-like Systems) – Maximum Efficiency
If you’re on macOS or Linux, the native find command is heavily optimized and will outperform any Python-based traversal for large directories. We can call it directly via subprocess to get the same speed as running find . -newermt in the shell.
import subprocess from datetime import datetime def find_files_newer_than_system(root_dir, target_date_str): # Build the find command cmd = [ 'find', root_dir, '-newermt', target_date_str, '-type', 'f' # Optional: limit to files (exclude directories) ] # Run the command and capture output try: result = subprocess.run( cmd, capture_output=True, text=True, check=True ) # Split output into individual file paths, filter out empty lines return [line.strip() for line in result.stdout.splitlines() if line.strip()] except subprocess.CalledProcessError as e: print(f"Error running find command: {e.stderr}") return [] # Example usage new_files = find_files_newer_than_system('.', '2018-01-17 03:28:46') for file in new_files: print(file)
Why this works:
- The system
finduses low-level filesystem APIs and is optimized for speed—no Python loop can match it for large directory trees. - It supports all the same flags as the shell command, so you can easily tweak it (e.g., exclude symlinks, limit depth).
3. Windows Alternative: Use PowerShell
If you’re on Windows, you can achieve similar efficiency by calling PowerShell’s Get-ChildItem cmdlet, which is optimized for the Windows filesystem:
import subprocess from datetime import datetime def find_files_newer_than_windows(root_dir, target_date_str): target_date = datetime.strptime(target_date_str, '%Y-%m-%d %H:%M:%S') # Format date for PowerShell ps_date_str = target_date.strftime('%Y-%m-%d %H:%M:%S') cmd = [ 'powershell', '-Command', f"Get-ChildItem -Path '{root_dir}' -Recurse -File | Where-Object {{ $_.LastWriteTime -gt '{ps_date_str}' }} | Select-Object -ExpandProperty FullName" ] try: result = subprocess.run( cmd, capture_output=True, text=True, check=True ) return [line.strip() for line in result.stdout.splitlines() if line.strip()] except subprocess.CalledProcessError as e: print(f"Error running PowerShell command: {e.stderr}") return [] # Example usage new_files = find_files_newer_than_windows('.', '2018-01-17 03:28:46') for file in new_files: print(file)
Final Recommendations
- Unix-like systems: Go with the system
findcall for maximum speed, especially with large directories. - Cross-platform needs: Use
os.scandir()—it’s the fastest pure-Python option and works everywhere.
内容的提问来源于stack exchange,提问作者Daniel Naftalovich

