You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python解析约14GB的超大JSON文件?

处理超大JSON文件的高效方案

我太懂你这种痛点了——14GB的JSON直接用json.load()全塞进内存,哪怕机器能扛住,也是对资源的浪费,而且扩展性极差。下面给你几个更优雅的解决方案:

1. 用流式解析库:ijson(最推荐)

ijson是专门为超大JSON设计的工具,它会逐块解析内容,完全不用把整个文件加载到内存里。

先安装库:

pip install ijson

如果你的JSON是一个大数组(比如[{}, {}, ...]),可以这样逐个处理元素:

import ijson

with open('myfile.json', 'r') as file:
    # 迭代数组里的每个对象,'item'是数组元素的通用路径
    for item in ijson.items(file, 'item'):
        # 在这里处理单个数据项,比如提取字段、做统计
        process_item(item)

要是JSON结构更复杂,比如嵌套的data.users数组,直接调整路径参数就行:ijson.items(file, 'data.users.item')。

2. 处理NDJSON格式(每行一个JSON对象)

如果你的超大JSON其实是每行一个独立的JSON对象(也就是NDJSON/JSON Lines格式),那处理起来更简单,完全不用额外库:

import json

with open('myfile.json', 'r') as file:
    for line in file:
        # 跳过空行(避免解析报错)
        if line.strip():
            item = json.loads(line)
            process_item(item)

这种方式内存占用极低,因为每次只加载一行内容。

3. 用内置json模块的raw_decode(无需额外依赖)

要是不想装第三方库,也可以用Python内置的json.JSONDecoder手动处理流——它会从输入里解析出一个完整JSON对象,同时返回下一个对象的起始位置:

import json

def stream_json(file_path):
    with open(file_path, 'r') as file:
        decoder = json.JSONDecoder()
        buffer = ''
        for line in file:
            buffer += line.strip()
            while buffer:
                try:
                    obj, idx = decoder.raw_decode(buffer)
                    yield obj
                    buffer = buffer[idx:]
                except ValueError:
                    # 缓冲区内容不够解析一个完整对象,继续读下一行
                    break

# 使用示例
for item in stream_json('myfile.json'):
    process_item(item)

这个方法适合不想引入外部依赖的场景,就是代码相对繁琐一点。


内容的提问来源于stack exchange,提问作者0xfede7c8

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:57:02