You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python修复无括号无逗号分隔的无效大JSON文件

修复格式错误的大体积JSON文件(Python实现)

你的JSON文件本质是JSON Lines格式(每行一个独立JSON对象),要转成标准JSON,只需给整体包裹数组括号,并在条目间添加逗号。针对大文件,不能一次性加载到内存,推荐逐行处理:

基础处理代码

直接完成格式转换,适合确定每行都是有效JSON对象的场景:

import json

input_file = "待修复的文件.json"
output_file = "修复后的文件.json"

with open(input_file, 'r', encoding='utf-8') as in_f, open(output_file, 'w', encoding='utf-8') as out_f:
    out_f.write('[')
    is_first_entry = True
    for line in in_f:
        cleaned_line = line.strip()
        if not cleaned_line:  # 跳过空行
            continue
        if not is_first_entry:
            out_f.write(',')
        out_f.write(cleaned_line)
        is_first_entry = False
    out_f.write(']')

带格式校验的版本

如果不确定每行是否都是有效JSON,可以加入校验步骤,跳过错误行并提示:

import json

input_file = "待修复的文件.json"
output_file = "修复后的文件.json"

with open(input_file, 'r', encoding='utf-8') as in_f, open(output_file, 'w', encoding='utf-8') as out_f:
    out_f.write('[')
    is_first_entry = True
    for line_num, line in enumerate(in_f, 1):
        cleaned_line = line.strip()
        if not cleaned_line:
            continue
        try:
            # 验证当前行是合法JSON
            json.loads(cleaned_line)
        except json.JSONDecodeError as err:
            print(f"第{line_num}行格式错误,已跳过: {err}")
            continue
        if not is_first_entry:
            out_f.write(',')
        out_f.write(cleaned_line)
        is_first_entry = False
    out_f.write(']')

注意事项

  • 两种方法都是逐行读写,内存占用极低,适合GB级别的大文件
  • 如果你的JSON条目存在跨行情况(比如一个对象占多行),需要先对文件进行预处理合并完整对象,再执行上述转换

内容的提问来源于stack exchange,提问作者José Carlos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 10:17:18