You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速大型JSON文件导入?现有串行导入速度过慢

提速大型JSON Lines文件导入的方法

你的代码目前串行处理每个JSON文件,且依赖标准库json解析、用动态变量管理结果,这些都可能拖慢速度,以下是几种可行的提速方案:

一、并行化IO处理

JSON文件读取属于IO密集型任务,用多进程并行处理多个文件,能大幅减少等待时间。推荐用concurrent.futures实现,代码示例:

import json
from concurrent.futures import ProcessPoolExecutor

def load_single_file(file_path):
    papers = []
    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            papers.append(json.loads(line.strip()))
    return papers

if __name__ == '__main__':
    file_list = [f'part_00{x}.json' for x in range(8)]
    # 进程池默认进程数等于CPU核心数,可通过max_workers指定
    with ProcessPoolExecutor() as executor:
        # all_papers是列表,每个元素对应一个文件的解析结果
        all_papers = list(executor.map(load_single_file, file_list))
    # 如需单独访问某文件结果,用all_papers[0]、all_papers[1]...即可

二、替换高性能JSON解析库

标准库json的解析速度不算顶尖,换成ujson或orjson这类专门优化的库,能直接提升解析效率:

  1. 安装依赖(二选一):
pip install ujson
# 或速度更快的orjson,部分系统需额外编译依赖
pip install orjson
  1. 修改解析代码(以ujson为例):
import ujson

def load_single_file(file_path):
    papers = []
    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            papers.append(ujson.loads(line.strip()))
    return papers

三、优化变量与IO操作

原代码用globals()动态创建变量的方式既不优雅也有额外开销,用列表统一存储结果能减少内存碎片、方便后续处理。此外,用二进制模式打开文件并解码,可降低IO层的转换开销:

def load_single_file(file_path):
    papers = []
    with open(file_path, 'rb') as f:
        for line in f:
            papers.append(ujson.loads(line.decode('utf-8').strip()))
    return papers

四、批量解析优化

如果文件是标准JSON Lines格式(每行一个JSON对象),可一次性读取所有行再批量解析,减少循环的微小开销:

def load_single_file(file_path):
    with open(file_path, 'rb') as f:
        lines = f.read().splitlines()
    return [ujson.loads(line.decode('utf-8')) for line in lines]

内容的提问来源于stack exchange,提问作者Lusian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 13:35:22