You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取大量小JSON文件:首次耗时久的原因及优化方案

问题解答

1. 首次读取耗时久、后续运行快的原因

  • 首次读取时,文件数据存储在物理硬盘(无论HDD还是SSD),需要从存储设备读取数据,物理IO的速度远低于内存读写速度。加上20000个小文件需要频繁执行文件打开、寻址操作,进一步增加了耗时。
  • 首次读取完成后,Windows系统会自动将这些文件的数据缓存到**内存页缓存(Page Cache)**中。后续再次读取时,系统直接从内存缓存中获取数据,无需再访问物理硬盘,因此耗时大幅缩短。

2. 缩短首次读取耗时的方法

(1)并行读取文件

利用多线程同时读取多个文件,减少串行等待的IO时间。示例代码:

from time import perf_counter
from glob import glob
from concurrent.futures import ThreadPoolExecutor

def load_file(file):
    with open(file, "r") as f:
        return f.read()

def load_all_jsons_parallel():
    t_i = perf_counter()
    json_files = glob("*json")
    # 线程数可根据磁盘IO能力调整,避免过多线程引发IO竞争
    with ThreadPoolExecutor(max_workers=16) as executor:
        executor.map(load_file, json_files)
    t_f = perf_counter()
    return t_f - t_i

print(load_all_jsons_parallel())

(2)合并小文件

将多个小JSON文件合并为少量大文件,减少文件打开、寻址的重复开销。比如按业务分类、时间区间等规则合并文件,后续读取时直接读取大文件再拆分数据。

(3)更换高速存储设备

如果当前使用机械硬盘(HDD),更换为固态硬盘(SSD)可大幅提升物理IO速度,从根源上降低首次读取的基础耗时。

(4)优化文件列表获取方式

用os.scandir()替代glob(),前者在遍历大量文件时性能更优,能减少获取文件列表的耗时:

from time import perf_counter
import os

def load_all_jsons_scandir():
    t_i = perf_counter()
    with os.scandir() as entries:
        for entry in entries:
            if entry.name.endswith(".json") and entry.is_file():
                with open(entry.path, "r") as f:
                    f.read()
    t_f = perf_counter()
    return t_f - t_i

print(load_all_jsons_scandir())

(5)预触发系统缓存

如果有提前准备的时间窗口,可先通过轻量IO操作(比如快速遍历文件并读取首字节)让系统将文件缓存到内存中,后续正式读取时直接调用缓存数据。


内容的提问来源于stack exchange,提问作者MKMS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 01:45:36