5GB JSON文件拆分性能优化:提速现有Python处理代码
5GB JSON文件分片优化需求
我有一个5GB的JSON文件,每行是一个独立的JSON对象,需要将其拆分为格式规范、分片大小可配置的多个JSON文件。当前代码因全量读取数据统计记录数,单文件处理耗时约18分钟,急需提速。
输入输出样例
输入样例
{"job": "developer"} {"job": "taxi driver"} {"job": "police"}
输出样例
{ "ROOT": [ { "job": "developer" }, { "job": "taxi driver" }, { "job": "police" } ] }
现有代码
import os import json import glob import time import shutil start = time.time() def filesplit(fInputname,farchivedir,foutdir): File_Extension = '.json' # using partition() # String till Substring x=foutdir.find(File_Extension) res=foutdir[0:x] with open(fInputname, 'r', encoding='utf-8') as f1: ll = [json.loads(line.strip()) for line in f1.readlines()] #Total number of records in the input json file print(len(ll)) #50000 means we getting splits of 50000 json objects size_of_the_split=50000 total = len(ll) // size_of_the_split #Number of files getting generated print(total+1) for i in range(total+1): jsonData=ll[i * size_of_the_split:(i + 1) * size_of_the_split] json.dump( {'ROOT': jsonData}, open(res + "_" + str(i+1) + ".json", 'w',encoding='utf8'), ensure_ascii=False, indent=True) shutil.move(fInputname,farchivedir) for name in glob.glob("C:\\Users\\JSON\\Input\\*.json"): print(name) filesplit(name, name.replace('C:\\Users\\JSON\\Input','C:\\Users\\JSON\\OriginalFiles_BKP'),name.replace('C:\\Users\\JSON\\Input','C:\\Users\\JSON\\Output')) end = time.time() print('completed') print("The time of execution of above program is :", (end-start) * 10**3, "ms")
优化方案
现有代码的核心问题是一次性读取全量数据到内存,既占用大量内存,又需要遍历两次(一次读取统计,一次分片写入)。优化思路改为逐行读取、边读边写,无需预先统计总行数,降低内存占用同时减少IO耗时。
优化后代码
import json import glob import time import shutil def filesplit(f_input, f_archive_dir, f_out_dir, split_size=50000): # 生成输出文件前缀 ext = '.json' prefix = f_out_dir[:f_out_dir.find(ext)] if ext in f_out_dir else f_out_dir file_index = 1 current_batch = [] with open(f_input, 'r', encoding='utf-8') as infile: for line in infile: line = line.strip() if not line: continue # 跳过空行 try: obj = json.loads(line) current_batch.append(obj) # 当批次达到指定大小,写入文件 if len(current_batch) == split_size: output_path = f"{prefix}_{file_index}.json" with open(output_path, 'w', encoding='utf-8') as outfile: json.dump({'ROOT': current_batch}, outfile, ensure_ascii=False, indent=True) print(f"已写入文件: {output_path}") current_batch = [] file_index += 1 except json.JSONDecodeError as e: print(f"解析JSON行失败: {e}, 行内容: {line}") continue # 处理剩余不足一个批次的数据 if current_batch: output_path = f"{prefix}_{file_index}.json" with open(output_path, 'w', encoding='utf-8') as outfile: json.dump({'ROOT': current_batch}, outfile, ensure_ascii=False, indent=True) print(f"已写入文件: {output_path}") # 归档原文件 shutil.move(f_input, f_archive_dir) print(f"原文件已归档至: {f_archive_dir}") if __name__ == "__main__": start = time.time() # 遍历输入目录下的JSON文件 for input_path in glob.glob("C:\\Users\\JSON\\Input\\*.json"): print(f"开始处理文件: {input_path}") archive_path = input_path.replace('C:\\Users\\JSON\\Input', 'C:\\Users\\JSON\\OriginalFiles_BKP') output_path = input_path.replace('C:\\Users\\JSON\\Input', 'C:\\Users\\JSON\\Output') filesplit(input_path, archive_path, output_path, split_size=50000) end = time.time() print("处理完成") print(f"总执行时间: {(end - start) * 10**3:.2f} ms")
优化点说明
- 逐行读取处理:避免一次性加载5GB数据到内存,内存占用仅为当前批次的大小(约50000个JSON对象),大幅降低内存压力。
- 边读边写:无需预先统计总行数,读取过程中达到批次大小就立即写入文件,减少一次全量遍历的IO耗时。
- 错误处理:增加JSON解析错误捕获,避免单个无效行导致整个程序崩溃。
- 代码结构优化:使用
if __name__ == "__main__"规范入口,参数更清晰,可读性更强。
内容的提问来源于stack exchange,提问作者sk1893
相关产品推荐
相关产品推荐

