You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何无需遍历全目录快速筛选指定日期之后的文件?

Efficiently Find Files Newer Than a Date in Python (No Unnecessary Traversal Overhead)

Great question—full directory traversals can be a total slog with large folders, right? The good news is we have options in Python that match (or come close to) the efficiency of find . -newermt on Unix-like systems. Let’s break down the best approaches:

1. Use os.scandir() (Python 3.5+) – Fast Native Python Approach

os.scandir() is a game-changer for directory traversal because it caches file metadata (like modification time) when it reads the directory, avoiding the extra os.stat() calls that slow down older methods like os.walk(). It still traverses the directory tree, but it’s the fastest pure-Python method available.

Here’s a recursive implementation that filters files newer than your target date:

import os
from datetime import datetime

def find_files_newer_than(root_dir, target_date_str):
    # Convert target date string to a Unix timestamp (seconds since epoch)
    target_timestamp = datetime.strptime(target_date_str, '%Y-%m-%d %H:%M:%S').timestamp()

    for entry in os.scandir(root_dir):
        if entry.is_file(follow_symlinks=False):
            # Access cached mtime directly from the DirEntry object
            if entry.stat().st_mtime > target_timestamp:
                yield entry.path
        elif entry.is_dir(follow_symlinks=False):
            # Recursively check subdirectories
            yield from find_files_newer_than(entry.path, target_date_str)

# Example usage
for file_path in find_files_newer_than('.', '2018-01-17 03:28:46'):
    print(file_path)

Why this works:

  • os.scandir() returns DirEntry objects that already hold stat info, so we don’t waste time re-querying the filesystem for each file.
  • It’s cross-platform (works on Windows, macOS, Linux) and doesn’t rely on external tools.

2. Call the System’s find Command (Unix-like Systems) – Maximum Efficiency

If you’re on macOS or Linux, the native find command is heavily optimized and will outperform any Python-based traversal for large directories. We can call it directly via subprocess to get the same speed as running find . -newermt in the shell.

import subprocess
from datetime import datetime

def find_files_newer_than_system(root_dir, target_date_str):
    # Build the find command
    cmd = [
        'find', root_dir,
        '-newermt', target_date_str,
        '-type', 'f'  # Optional: limit to files (exclude directories)
    ]

    # Run the command and capture output
    try:
        result = subprocess.run(
            cmd,
            capture_output=True,
            text=True,
            check=True
        )
        # Split output into individual file paths, filter out empty lines
        return [line.strip() for line in result.stdout.splitlines() if line.strip()]
    except subprocess.CalledProcessError as e:
        print(f"Error running find command: {e.stderr}")
        return []

# Example usage
new_files = find_files_newer_than_system('.', '2018-01-17 03:28:46')
for file in new_files:
    print(file)

Why this works:

  • The system find uses low-level filesystem APIs and is optimized for speed—no Python loop can match it for large directory trees.
  • It supports all the same flags as the shell command, so you can easily tweak it (e.g., exclude symlinks, limit depth).

3. Windows Alternative: Use PowerShell

If you’re on Windows, you can achieve similar efficiency by calling PowerShell’s Get-ChildItem cmdlet, which is optimized for the Windows filesystem:

import subprocess
from datetime import datetime

def find_files_newer_than_windows(root_dir, target_date_str):
    target_date = datetime.strptime(target_date_str, '%Y-%m-%d %H:%M:%S')
    # Format date for PowerShell
    ps_date_str = target_date.strftime('%Y-%m-%d %H:%M:%S')

    cmd = [
        'powershell',
        '-Command',
        f"Get-ChildItem -Path '{root_dir}' -Recurse -File | Where-Object {{ $_.LastWriteTime -gt '{ps_date_str}' }} | Select-Object -ExpandProperty FullName"
    ]

    try:
        result = subprocess.run(
            cmd,
            capture_output=True,
            text=True,
            check=True
        )
        return [line.strip() for line in result.stdout.splitlines() if line.strip()]
    except subprocess.CalledProcessError as e:
        print(f"Error running PowerShell command: {e.stderr}")
        return []

# Example usage
new_files = find_files_newer_than_windows('.', '2018-01-17 03:28:46')
for file in new_files:
    print(file)

Final Recommendations

  • Unix-like systems: Go with the system find call for maximum speed, especially with large directories.
  • Cross-platform needs: Use os.scandir()—it’s the fastest pure-Python option and works everywhere.

内容的提问来源于stack exchange,提问作者Daniel Naftalovich

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:18:06