如何高效获取目录中最旧的N个文件名?(Python/Shell)
First, let's directly answer your question about the ls -Art | head -n 1000 command:
- Does it work? Yes, technically. The
-tflag sorts files by modification time (newest first),-rreverses that order to oldest first, and-Askips the.and..directories. Piping tohead -n 1000will give you the first 1000 oldest files. - Does it list all files first? Unfortunately, yes.
lshas to read every directory entry, sort all of them, and then send the output tohead. For 150k files, this still involves processing all entries, which can be slow—maybe not as bad as your glob approach, but not optimal. Also, parsinglsoutput is risky if your filenames contain spaces, newlines, or special characters, sincelsmight mangle those.
Better Linux Shell Solutions
Option 1: Use find with sorting (more reliable)
This approach avoids the pitfalls of ls and is more efficient for large directories. It retrieves modification times as epoch timestamps, sorts numerically, and extracts the oldest 1000 files:
find /path/to/your/folder -maxdepth 1 -type f -name "*.json" -printf "%T@ %p\n" | sort -n | head -n 1000 | awk '{print $2}'
Breakdown:
-maxdepth 1: Only look in the target directory (no subfolders)-type f: Only include files (not directories)-name "*.json": Filter for JSON files-printf "%T@ %p\n": Print epoch modification time + file pathsort -n: Sort numerically by epoch time (oldest first)head -n 1000: Keep only the first 1000 entriesawk '{print $2}': Extract the file path (ignoring the epoch timestamp)
This handles special filenames correctly and is faster than ls for large datasets.
Python Solutions
Your original glob approach is slow because it retrieves all filenames first, then you'd have to stat each one to get modification times (which adds extra system calls). Here are two better approaches:
Option 1: Use os.scandir() with a heap (memory-efficient)
os.scandir() is faster than glob or os.listdir() because it retrieves file attributes (like modification time) during directory traversal, avoiding extra stat calls. Using a max-heap lets us keep only the oldest 1000 files in memory, which is ideal for large directories:
import os import heapq def get_oldest_json_files(folder_path, count=1000): max_heap = [] with os.scandir(folder_path) as entries: for entry in entries: if entry.is_file() and entry.name.endswith('.json'): # Get modification time (seconds since epoch) mtime = entry.stat().st_mtime # Use negative mtime since heapq is a min-heap (we want max-heap behavior) if len(max_heap) < count: heapq.heappush(max_heap, (-mtime, entry.path)) else: # Replace the newest file in the heap if current is older if mtime < -max_heap[0][0]: heapq.heappop(max_heap) heapq.heappush(max_heap, (-mtime, entry.path)) # Sort the heap results to get oldest first oldest_files = [path for neg_mtime, path in sorted(max_heap, reverse=True)] return oldest_files
Option 2: Use os.scandir() with full sort (simpler, uses more memory)
If memory isn't a constraint, you can collect all JSON files with their modification times, sort, and take the first 1000. This is simpler than the heap approach but stores all 150k entries temporarily:
import os def get_oldest_json_files(folder_path, count=1000): json_files = [] with os.scandir(folder_path) as entries: for entry in entries: if entry.is_file() and entry.name.endswith('.json'): mtime = entry.stat().st_mtime json_files.append((mtime, entry.path)) # Sort by modification time (oldest first) json_files.sort() # Extract the paths of the first 'count' files return [path for mtime, path in json_files[:count]]
Both Python approaches are significantly faster than your original glob method because they avoid redundant system calls to fetch file metadata.
内容的提问来源于stack exchange,提问作者kubera kalyan

