You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python将JSON Lines格式文件转换为JSON数组?

嘿,这个问题我太熟悉了!处理6GB这种超大体积的JSON行文件转成JSON数组,核心原则就是绝对不能把整个文件一次性读进内存——不然分分钟爆内存报错。下面给你两种实用的Python实现方案,按需选就行:

方案一:原生Python实现(无需额外依赖)

这种方法完全用Python标准库,不用装任何第三方包,内存占用极低,适合所有场景:

import json

input_path = "你的大文件路径.jsonl"
output_path = "输出的JSON数组文件.json"

with open(output_path, 'w', encoding='utf-8') as out_file:
    # 先写入JSON数组的左括号
    out_file.write('[')
    first_line = True
    
    with open(input_path, 'r', encoding='utf-8') as in_file:
        for line in in_file:
            # 去掉行首尾的空白字符,跳过空行
            cleaned_line = line.strip()
            if not cleaned_line:
                continue
            
            # 可选:验证每行的JSON格式,跳过无效行
            try:
                json.loads(cleaned_line)
            except json.JSONDecodeError as e:
                print(f"跳过无效JSON行: {cleaned_line[:50]}... 错误原因: {e}")
                continue
            
            # 除了第一行,其他行前面要加逗号分隔
            if not first_line:
                out_file.write(',')
            out_file.write(cleaned_line)
            first_line = False
            
            # 可选:强制刷新缓冲区,避免数据积压
            out_file.flush()
    
    # 最后写入JSON数组的右括号
    out_file.write(']')

这个方案的优点:

  • 零依赖,直接运行
  • 内存占用几乎可以忽略,只保留当前处理的一行数据
  • 自带错误处理,不会因为某一行格式错误导致整个任务失败

方案二:用jsonlines库简化代码(推荐经常处理这类文件的场景)

如果你经常和JSON Lines格式的文件打交道,装个jsonlines库能让代码更简洁,它会自动帮你处理行解析、空行跳过等细节:

首先先安装库:

pip install jsonlines

然后是代码:

import json
import jsonlines

input_path = "你的大文件路径.jsonl"
output_path = "输出的JSON数组文件.json"

with open(output_path, 'w', encoding='utf-8') as out_file:
    out_file.write('[')
    first_item = True
    
    # jsonlines会自动迭代每行的JSON对象
    with jsonlines.open(input_path, mode='r') as reader:
        for obj in reader:
            if not first_item:
                out_file.write(',')
            # 把Python对象序列化为JSON字符串写入
            json.dump(obj, out_file)
            first_item = False
            out_file.flush()
    
    out_file.write(']')

额外注意事项:

  1. 编码问题:始终指定encoding='utf-8',避免不同系统下的乱码问题
  2. 性能优化:如果文件读写速度慢,可以给open函数加buffering=1024*1024参数(设置1MB缓冲区),提升IO效率
  3. 验证结果:转换完成后,可以用命令行工具jq快速验证格式:jq '.' output_path.json,或者用Python加载一小部分内容检查

内容的提问来源于stack exchange,提问作者yokjc232

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 03:33:10