如何从超大目录中高效实现无放回随机懒加载文件?
百万级文件无放回随机采样的高效实现方案
针对目录中百万级且持续增长的文件,以下是几种替代os.listdir()的高效采样方案,适配不同场景需求:
1. 用os.scandir()替代os.listdir()
os.scandir()是Python 3.5+引入的目录遍历接口,性能远优于os.listdir()——它直接返回包含文件元信息的DirEntry对象,避免了额外的系统调用,且支持批量读取目录项,在大目录下速度提升明显。
import os import random def sample_files_with_scandir(path, sample_size): all_files = [] # 遍历目录,仅收集文件(排除子目录) with os.scandir(path) as entries: for entry in entries: if entry.is_file(): all_files.append(entry.name) # 无放回随机采样 return random.sample(all_files, sample_size)
适用场景:需要一次性获取全量文件列表后采样,内存足以容纳所有文件名的情况。
2. 蓄水池采样(内存友好+动态增长适配)
如果文件数量持续增长且内存有限,蓄水池采样算法无需加载全量文件列表,仅维护一个固定大小的“蓄水池”,遍历过程中动态更新采样结果,保证每个文件被选中的概率均等。
import os import random def reservoir_sample(path, sample_size): reservoir = [] file_count = 0 with os.scandir(path) as entries: for entry in entries: if entry.is_file(): file_count += 1 # 蓄水池未满时直接加入 if len(reservoir) < sample_size: reservoir.append(entry.name) else: # 生成随机索引,小于采样大小则替换蓄水池中的元素 rand_idx = random.randint(0, file_count - 1) if rand_idx < sample_size: reservoir[rand_idx] = entry.name return reservoir
适用场景:文件数量极大、内存紧张,或需要适配持续增长的目录(每次运行自动采样当前所有文件)。
3. 预维护文件索引(高频采样最优解)
若需要频繁执行采样操作,每次遍历目录的开销会累积,此时可以用轻量数据库(如SQLite)或文本文件维护文件索引,定期增量更新索引,采样时直接从索引中读取。
初始化索引
import sqlite3 import os def init_file_index(db_path, dir_path): conn = sqlite3.connect(db_path) cursor = conn.cursor() # 创建唯一索引避免重复记录 cursor.execute('CREATE TABLE IF NOT EXISTS files (filename TEXT UNIQUE)') # 批量插入现有文件 with os.scandir(dir_path) as entries: file_list = [(entry.name,) for entry in entries if entry.is_file()] cursor.executemany('INSERT OR IGNORE INTO files VALUES (?)', file_list) conn.commit() conn.close()
增量更新索引(定期执行或文件新增时触发)
def update_file_index(db_path, dir_path): conn = sqlite3.connect(db_path) cursor = conn.cursor() # 获取已记录的文件名集合 cursor.execute('SELECT filename FROM files') existing_files = set(row[0] for row in cursor.fetchall()) # 收集新增文件并插入 new_files = [] with os.scandir(dir_path) as entries: for entry in entries: if entry.is_file() and entry.name not in existing_files: new_files.append((entry.name,)) cursor.executemany('INSERT INTO files VALUES (?)', new_files) conn.commit() conn.close()
从索引采样
def sample_from_index(db_path, sample_size): conn = sqlite3.connect(db_path) cursor = conn.cursor() # 利用SQLite内置的随机排序实现无放回采样 cursor.execute(f'SELECT filename FROM files ORDER BY RANDOM() LIMIT {sample_size}') sample = [row[0] for row in cursor.fetchall()] conn.close() return sample
适用场景:高频采样需求,目录文件增长频率可预测(如定时新增)。
4. 系统原生工具调用(Linux/macOS极速方案)
在Linux或macOS系统上,利用原生命令行工具的性能优势,直接通过ls -U(不排序快速输出文件)和shuf(随机采样)完成操作,速度远快于Python遍历。
import subprocess def sample_with_system_tools(path, sample_size): # 过滤隐藏文件,采样指定数量的文件 cmd = f'ls -U "{path}" | grep -v "^\\." | shuf -n {sample_size}' result = subprocess.run( cmd, shell=True, capture_output=True, text=True, check=True ) # 按换行分割结果,过滤空行 return [fname for fname in result.stdout.strip().split('\n') if fname]
注意:需确保文件名不含换行符(机器学习数据集通常满足此条件);ls -U的兼容性需确认(主流Linux/macOS均支持)。
内容的提问来源于stack exchange,提问作者postnubilaphoebus
相关产品推荐
相关产品推荐

