Python读取大量小JSON文件:首次耗时久的原因及优化方案
问题解答
1. 首次读取耗时久、后续运行快的原因
- 首次读取时,文件数据存储在物理硬盘(无论HDD还是SSD),需要从存储设备读取数据,物理IO的速度远低于内存读写速度。加上20000个小文件需要频繁执行文件打开、寻址操作,进一步增加了耗时。
- 首次读取完成后,Windows系统会自动将这些文件的数据缓存到**内存页缓存(Page Cache)**中。后续再次读取时,系统直接从内存缓存中获取数据,无需再访问物理硬盘,因此耗时大幅缩短。
2. 缩短首次读取耗时的方法
(1)并行读取文件
利用多线程同时读取多个文件,减少串行等待的IO时间。示例代码:
from time import perf_counter from glob import glob from concurrent.futures import ThreadPoolExecutor def load_file(file): with open(file, "r") as f: return f.read() def load_all_jsons_parallel(): t_i = perf_counter() json_files = glob("*json") # 线程数可根据磁盘IO能力调整,避免过多线程引发IO竞争 with ThreadPoolExecutor(max_workers=16) as executor: executor.map(load_file, json_files) t_f = perf_counter() return t_f - t_i print(load_all_jsons_parallel())
(2)合并小文件
将多个小JSON文件合并为少量大文件,减少文件打开、寻址的重复开销。比如按业务分类、时间区间等规则合并文件,后续读取时直接读取大文件再拆分数据。
(3)更换高速存储设备
如果当前使用机械硬盘(HDD),更换为固态硬盘(SSD)可大幅提升物理IO速度,从根源上降低首次读取的基础耗时。
(4)优化文件列表获取方式
用os.scandir()替代glob(),前者在遍历大量文件时性能更优,能减少获取文件列表的耗时:
from time import perf_counter import os def load_all_jsons_scandir(): t_i = perf_counter() with os.scandir() as entries: for entry in entries: if entry.name.endswith(".json") and entry.is_file(): with open(entry.path, "r") as f: f.read() t_f = perf_counter() return t_f - t_i print(load_all_jsons_scandir())
(5)预触发系统缓存
如果有提前准备的时间窗口,可先通过轻量IO操作(比如快速遍历文件并读取首字节)让系统将文件缓存到内存中,后续正式读取时直接调用缓存数据。
内容的提问来源于stack exchange,提问作者MKMS
相关产品推荐
相关产品推荐

