如何将1330个文件按每10个为一组拆分,每组处理后输出1个结果文件
高效批量分组处理文件实现方案
你可以直接用Python列表切片+生成器的方式实现需求,代码简洁且运行效率高,具体实现如下:
基础实现(适合普通文件量场景)
直接通过列表切片完成定长分组,无需手动处理索引边界,最后不足10个文件的分组会自动保留:
import glob # 加载目标文件列表 file_list = glob.glob("你的文件路径匹配规则/*") # 每10个分为一组 batch_size = 10 file_batches = [file_list[i:i+batch_size] for i in range(0, len(file_list), batch_size)] # 逐组处理并输出结果 for batch_id, batch_files in enumerate(file_batches, start=1): # 替换为你自己的组内文件处理逻辑 batch_result = your_process_function(batch_files) # 输出对应组的结果文件 with open(f"输出路径/result_{batch_id}.txt", "w", encoding="utf-8") as f: f.write(batch_result)
大文件量优化实现(内存友好)
如果文件总量极大,可改用生成器实现,不会一次性加载全部分组到内存:
import glob def get_batches(file_list, batch_size=10): for idx in range(0, len(file_list), batch_size): yield file_list[idx:idx+batch_size] if __name__ == "__main__": file_list = glob.glob("你的文件路径匹配规则/*") for batch_id, batch_files in enumerate(get_batches(file_list), 1): # 组处理+输出逻辑同上 pass
并行加速实现(适合IO密集型处理场景)
如果组内文件处理是IO密集型操作,可以用线程池并行处理多个分组,大幅缩短耗时:
from concurrent.futures import ThreadPoolExecutor # 单组处理+保存逻辑封装 def process_single_batch(batch_info): batch_id, batch_files = batch_info batch_result = your_process_function(batch_files) with open(f"输出路径/result_{batch_id}.txt", "w", encoding="utf-8") as f: f.write(batch_result) if __name__ == "__main__": file_list = glob.glob("你的文件路径匹配规则/*") batch_size = 10 file_batches = [file_list[i:i+batch_size] for i in range(0, len(file_list), batch_size)] # 自定义最大并行线程数,IO密集型场景可以设为CPU核心数的2~4倍 with ThreadPoolExecutor(max_workers=8) as pool: pool.map(process_single_batch, enumerate(file_batches, 1))
内容的提问来源于stack exchange,提问作者Wind Circulation
相关产品推荐
相关产品推荐

