处理38GB大JSON文件时触发MemoryError的解决方案咨询
解决大JSON文件拆分时的MemoryError问题
你的代码触发MemoryError的核心原因是一次性将38GB的文件内容全部加载到内存列表ll中,这远远超出了常规内存的承载上限。以下是针对性的解决方法:
方法一:逐行流式处理(适用于每行一个独立JSON对象的文件)
如果你的JSON文件格式是每行一个独立的JSON对象(比如日志式JSON),可以逐行读取、分批写入,全程只在内存中保留当前批次的内容,内存占用极低:
import os import json input_path = os.path.join('folder/data.json') output_dir = 'result' os.makedirs(output_dir, exist_ok=True) # 确保输出目录存在 size_of_split = 1000000 # 每个拆分文件的条目数量 current_batch = [] file_index = 1 with open(input_path, 'r', encoding='utf-8') as f_in: for line in f_in: line = line.strip() if not line: continue # 跳过空行 try: obj = json.loads(line) current_batch.append(obj) # 当批次达到指定大小,写入文件 if len(current_batch) == size_of_split: output_path = os.path.join(output_dir, f'data_split{file_index}.json') with open(output_path, 'w', encoding='utf8') as f_out: json.dump(current_batch, f_out, ensure_ascii=False, indent=True) print(f"已完成文件: {output_path}") current_batch = [] file_index += 1 except json.JSONDecodeError as e: print(f"解析行出错: {e},跳过该行") # 处理剩余的不足一个批次的条目 if current_batch: output_path = os.path.join(output_dir, f'data_split{file_index}.json') with open(output_path, 'w', encoding='utf8') as f_out: json.dump(current_batch, f_out, ensure_ascii=False, indent=True) print(f"已完成最后一个文件: {output_path}")
方法二:流式解析JSON数组(适用于整个文件是大数组的情况)
如果你的JSON文件是一个完整的大数组(格式为[{}, {}, ...]),需要用流式解析库ijson来避免一次性加载整个数组:
- 先安装依赖:
pip install ijson
- 执行代码:
import os import json import ijson input_path = os.path.join('folder/data.json') output_dir = 'result' os.makedirs(output_dir, exist_ok=True) size_of_split = 1000000 current_batch = [] file_index = 1 with open(input_path, 'r', encoding='utf-8') as f_in: # 流式解析数组中的每个元素 for obj in ijson.items(f_in, 'item'): current_batch.append(obj) if len(current_batch) == size_of_split: output_path = os.path.join(output_dir, f'data_split{file_index}.json') with open(output_path, 'w', encoding='utf8') as f_out: json.dump(current_batch, f_out, ensure_ascii=False, indent=True) print(f"已完成文件: {output_path}") current_batch = [] file_index += 1 # 处理剩余条目 if current_batch: output_path = os.path.join(output_dir, f'data_split{file_index}.json') with open(output_path, 'w', encoding='utf8') as f_out: json.dump(current_batch, f_out, ensure_ascii=False, indent=True) print(f"已完成最后一个文件: {output_path}")
额外优化建议
- 如果不需要格式化输出(
indent=True),可以去掉该参数,能大幅提升写入速度并减小文件体积 - 若磁盘IO性能有限,可适当调小
size_of_split的值,避免单批次写入压力过大
内容的提问来源于stack exchange,提问作者Fatima
相关产品推荐
相关产品推荐

