You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将1330个文件按每10个为一组拆分,每组处理后输出1个结果文件

高效批量分组处理文件实现方案

你可以直接用Python列表切片+生成器的方式实现需求,代码简洁且运行效率高,具体实现如下:

基础实现(适合普通文件量场景)

直接通过列表切片完成定长分组,无需手动处理索引边界,最后不足10个文件的分组会自动保留:

import glob

# 加载目标文件列表
file_list = glob.glob("你的文件路径匹配规则/*")
# 每10个分为一组
batch_size = 10
file_batches = [file_list[i:i+batch_size] for i in range(0, len(file_list), batch_size)]

# 逐组处理并输出结果
for batch_id, batch_files in enumerate(file_batches, start=1):
    # 替换为你自己的组内文件处理逻辑
    batch_result = your_process_function(batch_files)
    # 输出对应组的结果文件
    with open(f"输出路径/result_{batch_id}.txt", "w", encoding="utf-8") as f:
        f.write(batch_result)

大文件量优化实现(内存友好)

如果文件总量极大,可改用生成器实现,不会一次性加载全部分组到内存:

import glob

def get_batches(file_list, batch_size=10):
    for idx in range(0, len(file_list), batch_size):
        yield file_list[idx:idx+batch_size]

if __name__ == "__main__":
    file_list = glob.glob("你的文件路径匹配规则/*")
    for batch_id, batch_files in enumerate(get_batches(file_list), 1):
        # 组处理+输出逻辑同上
        pass

并行加速实现(适合IO密集型处理场景)

如果组内文件处理是IO密集型操作,可以用线程池并行处理多个分组,大幅缩短耗时:

from concurrent.futures import ThreadPoolExecutor

# 单组处理+保存逻辑封装
def process_single_batch(batch_info):
    batch_id, batch_files = batch_info
    batch_result = your_process_function(batch_files)
    with open(f"输出路径/result_{batch_id}.txt", "w", encoding="utf-8") as f:
        f.write(batch_result)

if __name__ == "__main__":
    file_list = glob.glob("你的文件路径匹配规则/*")
    batch_size = 10
    file_batches = [file_list[i:i+batch_size] for i in range(0, len(file_list), batch_size)]
    # 自定义最大并行线程数,IO密集型场景可以设为CPU核心数的2~4倍
    with ThreadPoolExecutor(max_workers=8) as pool:
        pool.map(process_single_batch, enumerate(file_batches, 1))

内容的提问来源于stack exchange,提问作者Wind Circulation

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 15:12:03