如何使用Python将JSON Lines格式文件转换为JSON数组?
嘿,这个问题我太熟悉了!处理6GB这种超大体积的JSON行文件转成JSON数组,核心原则就是绝对不能把整个文件一次性读进内存——不然分分钟爆内存报错。下面给你两种实用的Python实现方案,按需选就行:
方案一:原生Python实现(无需额外依赖)
这种方法完全用Python标准库,不用装任何第三方包,内存占用极低,适合所有场景:
import json input_path = "你的大文件路径.jsonl" output_path = "输出的JSON数组文件.json" with open(output_path, 'w', encoding='utf-8') as out_file: # 先写入JSON数组的左括号 out_file.write('[') first_line = True with open(input_path, 'r', encoding='utf-8') as in_file: for line in in_file: # 去掉行首尾的空白字符,跳过空行 cleaned_line = line.strip() if not cleaned_line: continue # 可选:验证每行的JSON格式,跳过无效行 try: json.loads(cleaned_line) except json.JSONDecodeError as e: print(f"跳过无效JSON行: {cleaned_line[:50]}... 错误原因: {e}") continue # 除了第一行,其他行前面要加逗号分隔 if not first_line: out_file.write(',') out_file.write(cleaned_line) first_line = False # 可选:强制刷新缓冲区,避免数据积压 out_file.flush() # 最后写入JSON数组的右括号 out_file.write(']')
这个方案的优点:
- 零依赖,直接运行
- 内存占用几乎可以忽略,只保留当前处理的一行数据
- 自带错误处理,不会因为某一行格式错误导致整个任务失败
方案二:用
jsonlines库简化代码(推荐经常处理这类文件的场景) 如果你经常和JSON Lines格式的文件打交道,装个jsonlines库能让代码更简洁,它会自动帮你处理行解析、空行跳过等细节:
首先先安装库:
pip install jsonlines
然后是代码:
import json import jsonlines input_path = "你的大文件路径.jsonl" output_path = "输出的JSON数组文件.json" with open(output_path, 'w', encoding='utf-8') as out_file: out_file.write('[') first_item = True # jsonlines会自动迭代每行的JSON对象 with jsonlines.open(input_path, mode='r') as reader: for obj in reader: if not first_item: out_file.write(',') # 把Python对象序列化为JSON字符串写入 json.dump(obj, out_file) first_item = False out_file.flush() out_file.write(']')
额外注意事项:
- 编码问题:始终指定
encoding='utf-8',避免不同系统下的乱码问题 - 性能优化:如果文件读写速度慢,可以给
open函数加buffering=1024*1024参数(设置1MB缓冲区),提升IO效率 - 验证结果:转换完成后,可以用命令行工具
jq快速验证格式:jq '.' output_path.json,或者用Python加载一小部分内容检查
内容的提问来源于stack exchange,提问作者yokjc232
相关产品推荐
相关产品推荐

