You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

处理38GB大JSON文件时触发MemoryError的解决方案咨询

解决大JSON文件拆分时的MemoryError问题

你的代码触发MemoryError的核心原因是一次性将38GB的文件内容全部加载到内存列表ll中,这远远超出了常规内存的承载上限。以下是针对性的解决方法:

方法一:逐行流式处理(适用于每行一个独立JSON对象的文件)

如果你的JSON文件格式是每行一个独立的JSON对象(比如日志式JSON),可以逐行读取、分批写入,全程只在内存中保留当前批次的内容,内存占用极低:

import os
import json

input_path = os.path.join('folder/data.json')
output_dir = 'result'
os.makedirs(output_dir, exist_ok=True)  # 确保输出目录存在

size_of_split = 1000000  # 每个拆分文件的条目数量
current_batch = []
file_index = 1

with open(input_path, 'r', encoding='utf-8') as f_in:
    for line in f_in:
        line = line.strip()
        if not line:
            continue  # 跳过空行
        
        try:
            obj = json.loads(line)
            current_batch.append(obj)
            
            # 当批次达到指定大小,写入文件
            if len(current_batch) == size_of_split:
                output_path = os.path.join(output_dir, f'data_split{file_index}.json')
                with open(output_path, 'w', encoding='utf8') as f_out:
                    json.dump(current_batch, f_out, ensure_ascii=False, indent=True)
                print(f"已完成文件: {output_path}")
                current_batch = []
                file_index += 1
        except json.JSONDecodeError as e:
            print(f"解析行出错: {e},跳过该行")

# 处理剩余的不足一个批次的条目
if current_batch:
    output_path = os.path.join(output_dir, f'data_split{file_index}.json')
    with open(output_path, 'w', encoding='utf8') as f_out:
        json.dump(current_batch, f_out, ensure_ascii=False, indent=True)
    print(f"已完成最后一个文件: {output_path}")

方法二:流式解析JSON数组(适用于整个文件是大数组的情况)

如果你的JSON文件是一个完整的大数组(格式为[{}, {}, ...]),需要用流式解析库ijson来避免一次性加载整个数组:

  1. 先安装依赖:
pip install ijson
  1. 执行代码:
import os
import json
import ijson

input_path = os.path.join('folder/data.json')
output_dir = 'result'
os.makedirs(output_dir, exist_ok=True)

size_of_split = 1000000
current_batch = []
file_index = 1

with open(input_path, 'r', encoding='utf-8') as f_in:
    # 流式解析数组中的每个元素
    for obj in ijson.items(f_in, 'item'):
        current_batch.append(obj)
        
        if len(current_batch) == size_of_split:
            output_path = os.path.join(output_dir, f'data_split{file_index}.json')
            with open(output_path, 'w', encoding='utf8') as f_out:
                json.dump(current_batch, f_out, ensure_ascii=False, indent=True)
            print(f"已完成文件: {output_path}")
            current_batch = []
            file_index += 1

# 处理剩余条目
if current_batch:
    output_path = os.path.join(output_dir, f'data_split{file_index}.json')
    with open(output_path, 'w', encoding='utf8') as f_out:
        json.dump(current_batch, f_out, ensure_ascii=False, indent=True)
    print(f"已完成最后一个文件: {output_path}")

额外优化建议

  • 如果不需要格式化输出(indent=True),可以去掉该参数,能大幅提升写入速度并减小文件体积
  • 若磁盘IO性能有限,可适当调小size_of_split的值,避免单批次写入压力过大

内容的提问来源于stack exchange,提问作者Fatima

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 01:40:41