You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效获取目录中最旧的N个文件名?(Python/Shell)

Efficiently Getting the Oldest 1000 JSON Files

First, let's directly answer your question about the ls -Art | head -n 1000 command:

  • Does it work? Yes, technically. The -t flag sorts files by modification time (newest first), -r reverses that order to oldest first, and -A skips the . and .. directories. Piping to head -n 1000 will give you the first 1000 oldest files.
  • Does it list all files first? Unfortunately, yes. ls has to read every directory entry, sort all of them, and then send the output to head. For 150k files, this still involves processing all entries, which can be slow—maybe not as bad as your glob approach, but not optimal. Also, parsing ls output is risky if your filenames contain spaces, newlines, or special characters, since ls might mangle those.

Better Linux Shell Solutions

Option 1: Use find with sorting (more reliable)

This approach avoids the pitfalls of ls and is more efficient for large directories. It retrieves modification times as epoch timestamps, sorts numerically, and extracts the oldest 1000 files:

find /path/to/your/folder -maxdepth 1 -type f -name "*.json" -printf "%T@ %p\n" | sort -n | head -n 1000 | awk '{print $2}'

Breakdown:

  • -maxdepth 1: Only look in the target directory (no subfolders)
  • -type f: Only include files (not directories)
  • -name "*.json": Filter for JSON files
  • -printf "%T@ %p\n": Print epoch modification time + file path
  • sort -n: Sort numerically by epoch time (oldest first)
  • head -n 1000: Keep only the first 1000 entries
  • awk '{print $2}': Extract the file path (ignoring the epoch timestamp)

This handles special filenames correctly and is faster than ls for large datasets.


Python Solutions

Your original glob approach is slow because it retrieves all filenames first, then you'd have to stat each one to get modification times (which adds extra system calls). Here are two better approaches:

Option 1: Use os.scandir() with a heap (memory-efficient)

os.scandir() is faster than glob or os.listdir() because it retrieves file attributes (like modification time) during directory traversal, avoiding extra stat calls. Using a max-heap lets us keep only the oldest 1000 files in memory, which is ideal for large directories:

import os
import heapq

def get_oldest_json_files(folder_path, count=1000):
    max_heap = []
    with os.scandir(folder_path) as entries:
        for entry in entries:
            if entry.is_file() and entry.name.endswith('.json'):
                # Get modification time (seconds since epoch)
                mtime = entry.stat().st_mtime
                # Use negative mtime since heapq is a min-heap (we want max-heap behavior)
                if len(max_heap) < count:
                    heapq.heappush(max_heap, (-mtime, entry.path))
                else:
                    # Replace the newest file in the heap if current is older
                    if mtime < -max_heap[0][0]:
                        heapq.heappop(max_heap)
                        heapq.heappush(max_heap, (-mtime, entry.path))
    # Sort the heap results to get oldest first
    oldest_files = [path for neg_mtime, path in sorted(max_heap, reverse=True)]
    return oldest_files

Option 2: Use os.scandir() with full sort (simpler, uses more memory)

If memory isn't a constraint, you can collect all JSON files with their modification times, sort, and take the first 1000. This is simpler than the heap approach but stores all 150k entries temporarily:

import os

def get_oldest_json_files(folder_path, count=1000):
    json_files = []
    with os.scandir(folder_path) as entries:
        for entry in entries:
            if entry.is_file() and entry.name.endswith('.json'):
                mtime = entry.stat().st_mtime
                json_files.append((mtime, entry.path))
    # Sort by modification time (oldest first)
    json_files.sort()
    # Extract the paths of the first 'count' files
    return [path for mtime, path in json_files[:count]]

Both Python approaches are significantly faster than your original glob method because they avoid redundant system calls to fetch file metadata.


内容的提问来源于stack exchange,提问作者kubera kalyan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 12:47:39