You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从超大目录中高效实现无放回随机懒加载文件?

百万级文件无放回随机采样的高效实现方案

针对目录中百万级且持续增长的文件,以下是几种替代os.listdir()的高效采样方案,适配不同场景需求:

1. 用os.scandir()替代os.listdir()

os.scandir()是Python 3.5+引入的目录遍历接口,性能远优于os.listdir()——它直接返回包含文件元信息的DirEntry对象,避免了额外的系统调用,且支持批量读取目录项,在大目录下速度提升明显。

import os
import random

def sample_files_with_scandir(path, sample_size):
    all_files = []
    # 遍历目录,仅收集文件(排除子目录)
    with os.scandir(path) as entries:
        for entry in entries:
            if entry.is_file():
                all_files.append(entry.name)
    # 无放回随机采样
    return random.sample(all_files, sample_size)

适用场景:需要一次性获取全量文件列表后采样,内存足以容纳所有文件名的情况。

2. 蓄水池采样(内存友好+动态增长适配)

如果文件数量持续增长且内存有限,蓄水池采样算法无需加载全量文件列表,仅维护一个固定大小的“蓄水池”,遍历过程中动态更新采样结果,保证每个文件被选中的概率均等。

import os
import random

def reservoir_sample(path, sample_size):
    reservoir = []
    file_count = 0
    with os.scandir(path) as entries:
        for entry in entries:
            if entry.is_file():
                file_count += 1
                # 蓄水池未满时直接加入
                if len(reservoir) < sample_size:
                    reservoir.append(entry.name)
                else:
                    # 生成随机索引,小于采样大小则替换蓄水池中的元素
                    rand_idx = random.randint(0, file_count - 1)
                    if rand_idx < sample_size:
                        reservoir[rand_idx] = entry.name
    return reservoir

适用场景:文件数量极大、内存紧张,或需要适配持续增长的目录(每次运行自动采样当前所有文件)。

3. 预维护文件索引(高频采样最优解)

若需要频繁执行采样操作,每次遍历目录的开销会累积,此时可以用轻量数据库(如SQLite)或文本文件维护文件索引,定期增量更新索引,采样时直接从索引中读取。

初始化索引

import sqlite3
import os

def init_file_index(db_path, dir_path):
    conn = sqlite3.connect(db_path)
    cursor = conn.cursor()
    # 创建唯一索引避免重复记录
    cursor.execute('CREATE TABLE IF NOT EXISTS files (filename TEXT UNIQUE)')
    # 批量插入现有文件
    with os.scandir(dir_path) as entries:
        file_list = [(entry.name,) for entry in entries if entry.is_file()]
        cursor.executemany('INSERT OR IGNORE INTO files VALUES (?)', file_list)
    conn.commit()
    conn.close()

增量更新索引(定期执行或文件新增时触发)

def update_file_index(db_path, dir_path):
    conn = sqlite3.connect(db_path)
    cursor = conn.cursor()
    # 获取已记录的文件名集合
    cursor.execute('SELECT filename FROM files')
    existing_files = set(row[0] for row in cursor.fetchall())
    # 收集新增文件并插入
    new_files = []
    with os.scandir(dir_path) as entries:
        for entry in entries:
            if entry.is_file() and entry.name not in existing_files:
                new_files.append((entry.name,))
    cursor.executemany('INSERT INTO files VALUES (?)', new_files)
    conn.commit()
    conn.close()

从索引采样

def sample_from_index(db_path, sample_size):
    conn = sqlite3.connect(db_path)
    cursor = conn.cursor()
    # 利用SQLite内置的随机排序实现无放回采样
    cursor.execute(f'SELECT filename FROM files ORDER BY RANDOM() LIMIT {sample_size}')
    sample = [row[0] for row in cursor.fetchall()]
    conn.close()
    return sample

适用场景:高频采样需求,目录文件增长频率可预测(如定时新增)。

4. 系统原生工具调用(Linux/macOS极速方案)

在Linux或macOS系统上,利用原生命令行工具的性能优势,直接通过ls -U(不排序快速输出文件)和shuf(随机采样)完成操作,速度远快于Python遍历。

import subprocess

def sample_with_system_tools(path, sample_size):
    # 过滤隐藏文件,采样指定数量的文件
    cmd = f'ls -U "{path}" | grep -v "^\\." | shuf -n {sample_size}'
    result = subprocess.run(
        cmd,
        shell=True,
        capture_output=True,
        text=True,
        check=True
    )
    # 按换行分割结果,过滤空行
    return [fname for fname in result.stdout.strip().split('\n') if fname]

注意:需确保文件名不含换行符(机器学习数据集通常满足此条件);ls -U的兼容性需确认(主流Linux/macOS均支持)。

内容的提问来源于stack exchange,提问作者postnubilaphoebus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 00:06:17